Interesting work on long-horizon research agents.
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
Interesting work on long-horizon research agents. PrimeScientists can decide where a research agent spends its budget. Achieves 10.3% more reward with 50.6% fewer research attempts, under the same budget. That is PrimeScientist against AutoResearch on 12 AI research tasks. They treat deciding where to spend a research agent's budget as part of the agent's job. PrimeScientist keeps an executable plan tree of competing research directions and their outcomes. An adaptive MCTS policy reads the experimental feedback and the remaining budget and chooses whether to explore a new direction or con
Posted by DAIR.AI (132.7k followers) 1 h ago · 8 likes · 1.7k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: academy.dair.ai.
More AI work like this
- What if a model could predict the next word using just 2 to 4 previous words? Meet this… — @HuggingModels
- We are now using Claude Opus 5.5 to create blog-style overviews for every arXiv paper — @askalphaxiv
- Meet this text-generation transformer with a twist: it's a hard 3-gram model. Trained on… — @HuggingModels
- Meet a tiny language model with a big name: hard_3gram_4_6_384_babylm_10m_seed43. It's a… — @HuggingModels
- VLAs? WAMs? Are language and video models the right foundations for robotics? — @DJiafei
- Great paper from Google and colleagues. — @omarsar0
- New podcast with @datagenproc of @EpochAIResearch digging into the open questions… — @natolambert
- OPENAI'S AI JUST SOLVED A MATH PROBLEM THAT WAS SUPPOSED TO BE IMPOSSIBLE FOR COMPUTERS. — @ArkkDaily
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.9k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 19:32 UTC. Full method.