Brood War Bench
Recorded: Sept. 19, 2026, 8 p.m.
| Original | Summarized |
Brood War Bench Brood War BenchBy Ben SwerdlowWhich model wins at Brood War?Key takeawaysNone of the models played beyond a beginner level.Codex Astra is the clear leader beating all other models consistently.Grok models are not smart enough to play Brood War yet.Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking.Powered By FreestyleLeaderboardRankSystemWinsLossesAPMCost / gameWin rate🥇Codex Astra / xhigh18012.6$10.54100.0%🥈Codex Astra / medium16217.2$15.1188.9%🥉Claude Fable15312.6$12.2483.3%4Codex Astra / low14425.7$21.0777.8%5Codex 5.6 Sol / medium13510.1$5.1272.2%6Codex 5.6 Sol / low12618.1$9.2366.7%7Claude Opus 512610.5$20.7866.7%8Codex 5.6 Sol / xhigh1178.0$3.2361.1%9Codex 5.6 Luna / low9923.8$0.4250.0%10Codex 5.6 Terra / xhigh9915.8$2.1050.0%11Codex 5.6 Terra / medium81010.5$3.1544.4%12Codex 5.6 Terra / low81048.3$4.6544.4%13Codex 5.6 Luna / xhigh7115.2$0.1638.9%14Claude Sonnet7116.2$8.9838.9%15Codex 5.6 Luna / medium61214.8$0.3033.3%16Grok 4.6 / xhigh2152.8$0.6611.1%17Grok 4.6 / medium1163.2$0.795.6%18Claude Haiku0160.3$0.340.0%19Grok 4.6 / low0164.2$1.260.0%Show allBrood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn't done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.What I observed01Codex found cheese before it found macroCodex's strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn't communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build.This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing. In games where I helped direct Codex, it was much better at planning those moments and getting its subagents to work together.The persistence was real. In G009, after losing its army and main base, Codex 5.6 Terra / medium lifted its last Command Center and moved it toward the opposite corner. It survived for another six minutes.Six Probes cross the mapA Probe first, then Zealots in dripsThe last Command Center runs02Grok spent the game between actionsGrok 4.6 frequently produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit.The actions it did take rarely developed into a working control loop. In G003, Grok / xhigh made three Marines and never reached the enemy base. In G002, Grok / medium made two Zealots and also never crossed the map. These looked less like bad strategies than failures to keep observing and acting.Forty-three minutes, no army03Fable earnestly tried to play the gameI found myself rooting for Claude Fable in more than a few games. Fable usually tried to build an economy and climb the tech tree instead of stopping at the first unit available. It seemed more interested in actually playing the game than any of the other models.In G007 it reached a Lair, Spire, and Mutalisks and won. In G027 it added a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives before winning. Ambition did not guarantee execution: in G036 Fable reached a Factory and Academy but Opus 5 overran it.Fable gets MutalisksFable keeps climbingThe build does not become an armyNo agent here played beyond beginner levelEven Astra and Fable were unable to build complex army's, defend simple attacks or play concrete strategies. A beginner playing photon rush would win every single one of these games.That said, watching the agents play made me more excited than I have been in a while. This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do. I look forward to watching them get there.How the games developedCompareModelsModel + effortHarnessesRaceAll racesTerranProtossZergGame timeFirst 10 minutesFirst 15 minutesFirst 20 minutesFull gamesChoose up to fiveCodex AstraClaude FableCodex 5.6 SolClaude Opus 5Codex 5.6 LunaCodex 5.6 TerraClaude SonnetGrok 4.6Claude Haiku5:00 / 15:00Technology investmentCompleted research + upgrade levels00.511.520:005:0010:0015:00Codex Astra0.3Claude Fable0.2Grok 4.60WorkersCompleted workers alive081624320:005:0010:0015:00Codex Astra14.7Claude Fable16.7Grok 4.67.4Army sizeCompleted army and support units0204060800:005:0010:0015:00Codex Astra8.3Claude Fable7.6Grok 4.62StructuresCompleted buildings, including add-ons071421280:005:0010:0015:00Codex Astra6.2Claude Fable7.4Grok 4.64.5Minerals in the bankUnspent minerals, not income05001,0001,5002,0000:005:0010:0015:00Codex Astra240.6Claude Fable248.6Grok 4.6536.6Gas in the bankUnspent gas, not income05001,0001,5002,0000:005:0010:0015:00Codex Astra151Claude Fable326.9Grok 4.6244.5Supply usedIncludes production in progress03060901200:005:0010:0015:00Codex Astra26.5Claude Fable26.6Grok 4.611.3When games endedShare of games ending per 5-minute windowCodex Astra: median 8:10. 0:00 to before 5:00: 3.9%; 5:00 to before 10:00: 54.9%; 10:00 to before 15:00: 31.4%; 15:00 to before 20:00: 5.9%; 25:00 to before 30:00: 2%; 30:00 to before 35:00: 2%. Claude Fable: median 10:37. 5:00 to before 10:00: 38.9%; 10:00 to before 15:00: 38.9%; 15:00 to before 20:00: 22.2%. Grok 4.6: median 9:15. 0:00 to before 5:00: 2%; 5:00 to before 10:00: 51%; 10:00 to before 15:00: 23.5%; 15:00 to before 20:00: 13.7%; 20:00 to before 25:00: 2%; 40:00 to before 45:00: 7.8%.Codex AstraMedian 8:1060%Codex Astra: 3.9% (2 games) ended from 0:00 to before 5:00Codex Astra: 54.9% (28 games) ended from 5:00 to before 10:00Codex Astra: 31.4% (16 games) ended from 10:00 to before 15:00Codex Astra: 5.9% (3 games) ended from 15:00 to before 20:00Codex Astra: 0% (0 games) ended from 20:00 to before 25:00Codex Astra: 2% (1 games) ended from 25:00 to before 30:00Codex Astra: 2% (1 games) ended from 30:00 to before 35:00Codex Astra: 0% (0 games) ended from 35:00 to before 40:00Codex Astra: 0% (0 games) ended from 40:00 to before 45:00Claude FableMedian 10:3760%Claude Fable: 0% (0 games) ended from 0:00 to before 5:00Claude Fable: 38.9% (7 games) ended from 5:00 to before 10:00Claude Fable: 38.9% (7 games) ended from 10:00 to before 15:00Claude Fable: 22.2% (4 games) ended from 15:00 to before 20:00Claude Fable: 0% (0 games) ended from 20:00 to before 25:00Claude Fable: 0% (0 games) ended from 25:00 to before 30:00Claude Fable: 0% (0 games) ended from 30:00 to before 35:00Claude Fable: 0% (0 games) ended from 35:00 to before 40:00Claude Fable: 0% (0 games) ended from 40:00 to before 45:00Grok 4.6Median 9:1560%Grok 4.6: 2% (1 games) ended from 0:00 to before 5:00Grok 4.6: 51% (26 games) ended from 5:00 to before 10:00Grok 4.6: 23.5% (12 games) ended from 10:00 to before 15:00Grok 4.6: 13.7% (7 games) ended from 15:00 to before 20:00Grok 4.6: 2% (1 games) ended from 20:00 to before 25:00Grok 4.6: 0% (0 games) ended from 25:00 to before 30:00Grok 4.6: 0% (0 games) ended from 30:00 to before 35:00Grok 4.6: 0% (0 games) ended from 35:00 to before 40:00Grok 4.6: 7.8% (4 games) ended from 40:00 to before 45:000:0015:0030:0045:00Time-series charts show means of recorded player-runs at each game time. Finished games drop out; missing samples are not filled. Models pool their effort settings. Units and buildings count only once completed; army excludes workers, Overlords, eggs, larvae, and ammunition.Win rate vs. costLog scaleLinearAverage cost per game, using the same prices as the leaderboard. Codex and Sonnet costs are token-based estimates.CodexClaudeGrokWin rate0%25%50%75%100%$0.1$0.5$1$5$10$20Cost per game (USD, log scale)How the benchmark ranWe built a round-robin matrix of model and effort configurations and had every configuration play every other. The harness ran those matchups in parallel across Freestyle VMs, saving game-engine data and both agents' harness logs for each match.Head-to-head matrixRead across a row. W is a win, L is a loss, and T is a match that reached the benchmark time limit.Open the full 19 × 19 matrixW win L loss T time limit G000pendingSystem123456789101112131415161718191Codex Astra / xhigh-WWWWWWWWWWWWWWWWWW2Codex Astra / mediumL-WWWWWWWWWWLWWWWWW3Codex Astra / lowLL-WWWWWWWWWLLWWWWW4Codex 5.6 Sol / xhighLLL-LWWWLWWWLLWWWWW5Codex 5.6 Sol / mediumLLLW-WWWLWWWLWWWWWW6Codex 5.6 Sol / lowLLLLL-WWWWWWWWLWWWW7Codex 5.6 Luna / xhighLLLLLL-WLWLWLLLWWWW8Codex 5.6 Luna / mediumLLLLLLL-LLLLLWWWWWW9Codex 5.6 Luna / lowLLLWWLWW-WLLLLLWWWW10Codex 5.6 Terra / xhighLLLLLLLWL-WWLWWWWWW11Codex 5.6 Terra / mediumLLLLLLWWWL-LLLWWWWW12Codex 5.6 Terra / lowLLLLLLLWWLW-LLWWWWW13Claude FableLWWWWLWWWWWW-LWWWWW14Claude Opus 5LLWWLLWLWLWWW-WWWWW15Claude SonnetLLLLLWWLWLLLLL-WWWW16Claude HaikuLLLLLLLLLLLLLLL-LTT17Grok 4.6 / xhighLLLLLLLLLLLLLLLW-WT18Grok 4.6 / mediumLLLLLLLLLLLLLLLTL-W19Grok 4.6 / lowLLLLLLLLLLLLLLLTTL-Play your own matchBring your agent and play Brood War with friends.Play Brood WarPowered By Freestyle |
Ben Swerdlow conducted a benchmark experiment on various AI models to assess their capabilities in playing the real-time strategy game Brood War, using agents controlled by different models. The initial phase of the experiment revealed that none of the models played beyond a beginner level, suggesting a need for further development before attempting complex strategy. Codex Astra emerged as the dominant model, consistently winning across various configurations. Conversely, the Grok models were found to be insufficiently smart for strategic gameplay at this stage. The analysis of agent behavior provided several critical insights into the limitations of the models. Codex demonstrated a recurring tendency to prioritize disruption, often utilizing probes to attack workers or buildings, which successfully forced opponents to expend time considering the threat. However, this approach resulted in weaker sustained production and a failure to coordinate its subagents effectively; units were often sent into attacks individually rather than building up to a critical mass for synchronized assaults. The persistence demonstrated by some agents, such as Codex 5.6 Terra / medium surviving intense pressure, highlighted the capacity for emergent strategic shifts. Claude Fable exhibited a greater commitment to gameplay, focusing on building economies and climbing the technology tree rather than simply achieving the first available unit. While Fable showed high ambition, results were inconsistent, indicating that ambition does not guarantee successful execution, as demonstrated by instances where Fable built necessary structures but failed to secure victory against superior opponents. Grok, despite extensive reasoning, demonstrated a gap between its reasoning capabilities and practical command execution. The data indicated that Grok frequently engaged in long reasoning sequences while issuing very few actual command batches, suggesting a failure to translate internal thought processes into coherent, actionable military maneuvers. The benchmark involved structuring a round-robin matrix comparing different models and effort settings against one another across various segments of the game, including the first ten, fifteen, and twenty minutes, as well as full game outcomes. This matrix demonstrated how agents performed across different phases of the match. The quantitative results, including win rates, cost per game, and a time-series analysis of when games concluded, further differentiated the agents. For example, Codex Astra maintained a high win rate, and the time-series data indicated a tendency for games involving Codex Astra to conclude in the middle phases of the recorded time, suggesting a balanced pace of engagement. Claude Fable also showed a median game duration, indicating a sustained involvement in the match. Overall, the experiment emphasized that while the models possess potential, deeper strategic understanding and coordinated execution remain areas requiring significant development. |