Impressive paper from Salesforce.
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
Impressive paper from Salesforce. It discusses the importance of good verifiers for RL environments. Only 35.8% of the environments in the cleanest public RL collection for terminal agents passed Salesforce AI Research's audit. More details below: With the budget held at 3.5K environments, River-8B averaged 19.4 across four terminal benchmarks, against 17.7 for RL on 3.5K environments sampled at random from the same collection. The audit found reward errors in both directions. Some environments give reward 1 for copying a leaked answer or passing a weak verifier without doing the task. Ot
Posted by DAIR.AI (132.7k followers) 1 h ago · 19 likes · 2.6k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: academy.dair.ai.
More AI work like this
- Contrastive Language Models — @_akhaliq
- Yukon: Make Research Open & Multiplayer Again. — @sreeramkannan
- Must-read paper from Google on self-improving agent harnesses. — @omarsar0
- people arguing that AI won't get access to wet labs — @NathanpmYoung
- Skild AI trained S1 through 140 years of self-play in Nvidia Isaac Sim, with no human… — @wallstengine
- All-in-One Multilingual Scene Text Recognition with ScriptMoE — @HuggingPapers
- Banger paper from Microsoft Research and colleagues. — @omarsar0
- AI JUST PUT AN ARTIFICIAL FLY COLONY INSIDE AN IPHONE. — @Frank_web33
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.7k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 02:41 UTC. Full method.