RSI, this time from Xiaomi
This is a AI post classified by Jev as AI research (a model release), kept by the AI Radar because it carries real work, not commentary.
RSI, this time from Xiaomi Xiaomi is openly scaling AI toward self improvement Xiaomi’s MiMo-V2.6 is its biggest reinforcement learning scaling effort yet, pushing RL across coding, general reasoning, visual tasks, automation, and cybersecurity. The training reportedly processes roughly 1,568 samples and 2.7–3.7 billion tokens per RL step, with context lengths up to 1 million tokens. Instead of training on a narrow benchmark, Xiaomi uses multiple complex environments and stronger agentic graders to improve long horizon reasoning while also pushing the model toward more efficient solutions.
Posted by Dr Singularity (76k followers) 1 h ago · 103 likes · 3.7k views · view the original post on X. Kept by the AI Radar as AI research.
More AI work like this
- 👋 AI researchers, come say hello on Zhihu! — @ZhihuFrontier
- We ran 18 AI systems from 6 companies through an exam on human aging biology 🧬 Not one… — @andrewaiginin
- "Self-Organizing Agent Teams Learn to Reason Together" — @askalphaxiv
- Tencent ARC releases GAE: a geometry-native latent space for 3D-consistent world… — @HuggingPapers
- "RRSI: Regularized Recursive Self-Improvement of Agent Harnesses" — @askalphaxiv
- Tracking NBA players from a live camera feed onto a computer vision court map in real… — @0x0SojalSec
- Meet a tiny text generator trained on BabyLM data. This 10M parameter model uses a… — @HuggingModels
- What if a model could predict the next word using just 2 to 4 previous words? Meet this… — @HuggingModels
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.3k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 13:36 UTC. Full method.