RL teaches models to work longer, but reasoning is dependent on domain-specific…
This is a AI post classified by Jev as AI infra & evals (a free resource), kept by the AI Radar because it carries real work, not commentary.
RL teaches models to work longer, but reasoning is dependent on domain-specific post-training. Baseten's Head of Model Training @oneill_c sat down with @dwarkesh_sp to explain horizon generalization and what's next at the frontier. Full episode here: https://www.youtube.com/watch?v=PrSf7IOYu-I
Posted by Baseten (19.7k followers) 1 h ago · 9 likes · 392 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Cyber attackers move fast, so defenders have to move faster 🚨 — @cerebras
- Yesterday marked one month anniversary of @DarkbloomAI moving to paid tier on OpenRouter. — @gajesh
- Actually wait, answer quality might have gone up. But 25% savings is definitely holding. — @mmastrac
- Your agent may have completed the request without an error, but it retrieved the wrong… — @Snowflake
- We’ve improved prompt caching in the API for GPT-6, helping agents run faster and cost… — @OpenAIDevs
- This was so much fun. — @jvr0x
- What makes an AI factory energy efficient? — @NVIDIAAIInfra
- $QCOM UNVEILS NEW 2NM SNAPDRAGON CHIPS WITH BIG ON-DEVICE AI PUSH — @wallstengine
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 22:57 UTC. Full method.