AI Radar
Support
LiveUpdated 2026-09-22 05:59 UTC

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement”

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL…

This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL scaling run so far, openly. And it came along with a technical report, outlining how they've scaled RL across compute, environments, and grading. tl;dr The run costs $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Each RL step uses 1,568 prompts x 16 rollouts, producing ~25K trajectories and 2.7-3.7B tokens, with contexts up to 1M tokens. New techniques include groupwise agentic grading, which compares successful trajecto

Posted by alphaXiv (56.1k followers) 1 h ago · 16 likes · 1k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: alphaXiv.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 41k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 05:59 UTC. Full method.