有人做了一个"星际争霸"AI 对战测试(Brood War Bench),让当下主流大模型互相打即时战略,结果所有模型的水平都没超过新手。

This is a AI post classified by Jev as AI agents (a model release), kept by the AI Radar because it carries real work, not commentary.
有人做了一个"星际争霸"AI 对战测试(Brood War Bench),让当下主流大模型互相打即时战略,结果所有模型的水平都没超过新手。 《星际争霸:母巢之战》是 1998 年的经典即时战略游戏,也是 AI 研究的老朋友,2019 年 DeepMind 的 AlphaStar 就曾在这款游戏里击败职业选手。但那是专门训练的强化学习 AI。这次测试不同,它让通用大模型以 AI 智能体的形式直接上手操作,看看它们能不能自己建基地、造兵、打仗。 作者 Ben Swerdlow 原本只是做了一个"只能通过智能体操控"的星际争霸版本,拿来和朋友玩。没想到几个几乎没玩过星际的朋友表现还不错——他们说自己就下了句"去进攻"的命令,智能体就自己造了支小部队冲过去了。这让他好奇:如果完全让 AI 自己玩,能打到什么水平? 答案是:很菜,但很有意思。 排名第一的 Codex Astra 打了 18 场全胜,但它最擅长的不是正面作战,而是"骚扰":派一个采矿的工人(Probe)跑到对方基地捣乱。这招对 AI 对手特别好使,因为对面的智能体看到一个工人来了,会花几十秒思考该怎么办,期间什么都不干。而在正经的经济发展和大规模作战方面,Codex 反而比较弱,经常造一两个兵就往对面扔,而不是攒够兵力再出击。 Claude Fable 排第三,胜率 83.3%,是所有参赛模型里最"像在认真打游戏"的
Posted by 宝玉 (250k followers) 1 h ago · 3 likes · 1.6k views · view the original post on X. Kept by the AI Radar as AI agents. Tools mentioned: Agent StarCraft.
More AI work like this
- agents can run for longer than ever. — @tokensandai
- we ask AI for answers all day — @vikktorrrre
- yooo hold up, okay, hold on. — @RileyRalmuto
- "Our AI agents have already done 17,000 conversations this year. And booked 600+… — @jasonlk
- Open sourced "Jev " that runs 50x faster own Jarvis running locally now on mac 1GB RAM.… — @0x0SojalSec
- Banger paper on self-evolving ontologies for agents. — @omarsar0
- this is literally f**king insane — @gippp69
- Hopped on a Virtual Claude workshop this morning with @AnthropicAI and @cinkotweets to… — @JJEnglert
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 39.8k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 19:14 UTC. Full method.