“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement”
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL scaling run so far, openly. And it came along with a technical report, outlining how they've scaled RL across compute, environments, and grading. tl;dr The run costs $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Each RL step uses 1,568 prompts x 16 rollouts, producing ~25K trajectories and 2.7-3.7B tokens, with contexts up to 1M tokens. New techniques include groupwise agentic grading, which compares successful trajecto
Posted by alphaXiv (56.1k followers) 1 h ago · 16 likes · 1k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: alphaXiv.
More AI work like this
- 很有意思的图啊,MiMo 是想一步一个脚印靠自己爬~ — @Fei2411
- Hoping someone can explain to me what's going on here. — @lu__jasper
- $10,000 went out to solvers last week. — @yukonresearch
- The absolutely beautiful free hands-on book "Physics-Based Deep Learning Book" by… — @burkov
- Every mature systems project eventually needs a teaching version. — @PyTorch
- update: it broke :( — @pentestduck
- Scientific discovery poses significant challenges for measuring an AI system, as… — @Apodex_AI
- a model trained for 24 days has now resolved more than 100 long-standing open problems… — @IterIntellectus
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 41k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 05:59 UTC. Full method.