Impressive paper showing the impact of a good harness.
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
Impressive paper showing the impact of a good harness. Improves Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks without any change to the model. The gain comes from the harness. GAVEL keeps an explicit graph world model holding object relations, action preconditions and effects, and probabilistic beliefs about where unobserved objects are. Before the robot executes an LLM-generated action, the graph predicts what that action would do. Violations get caught, and the ones whose fix follows directly from the world model get repaired without calling the model again. Only errors that
Posted by DAIR.AI (132.4k followers) 1 h ago · 34 likes · 2.4k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: academy.dair.ai.
More AI work like this
- SoupFold from KAIST: no single co-folding model wins everywhere, so learn simple… — @biogerontology
- We built AI that actually knows aging biology. Longevity-LLM (9B params) beats OpenAI… — @biogerontology
- We need more examples like this in the open-source RL ecosystem — @adithya_s_k
- "Verbalizing Subliminal Learning Effects Using Text Optimization" — @askalphaxiv
- “What Does Privileged Information Add to On-Policy Self-Distillation?” — @askalphaxiv
- “Score Centering Stabilizes Off-Policy Reinforcement Learning” — @askalphaxiv
- 过去,让机器人学会一个新任务,往往需要大量数据采集、轨迹标注和针对性的训练。 — @qingke_ai
- Banger paper from Google Cloud AI Research. — @dair_ai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.7k posts from 4.8k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 03:20 UTC. Full method.