Must-read paper from Google on self-improving agent harnesses.
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
Must-read paper from Google on self-improving agent harnesses. If you auto-optimize your agent's harness, your eval score can go up while the agent gets worse on real tasks. This paper shows how to prevent that. Of five harness-evolution methods compared on agentic workspace tasks, RRSI scored the lowest on the tasks it evolved against and highest on all three out-of-distribution benchmarks. Automated harness evolution proposes edits to prompts, control flow, tools and memory, keeps the ones that raise the score, and repeats. The authors show this overfits the training tasks. Meta-Harness
Posted by elvis (321k followers) 1 h ago · 32 likes · 2.6k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: academy.dair.ai.
More AI work like this
- people arguing that AI won't get access to wet labs — @NathanpmYoung
- Skild AI trained S1 through 140 years of self-play in Nvidia Isaac Sim, with no human… — @wallstengine
- All-in-One Multilingual Scene Text Recognition with ScriptMoE — @HuggingPapers
- Banger paper from Microsoft Research and colleagues. — @omarsar0
- AI JUST PUT AN ARTIFICIAL FLY COLONY INSIDE AN IPHONE. — @Frank_web33
- The name HySparse2 really makes me think of Hy (Hunyuan). I even catch myself reading it… — @sheriyuo
- Xiaomi is skipping a generation. V2.6 still uses MiMo Hybrid-SWA, a very 2025 design;… — @teortaxesTex
- MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. — @_LuoFuli
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.2k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 21:31 UTC. Full method.