AI Radar
Support
LiveUpdated 2026-09-19 21:10 UTC

首先注意同时在训练2个模型, mimo-v2.6-pro 和 flash, 这个无疑是进入后训练阶段了, 写着RL嘛.

1秒烧10刀! 小米正在直播模型训练! 首先注意同时在训练2个模型, mimo-v2.6-pro 和 flash, 这个无疑是进入后训练阶段了, 写着RL嘛. 然后注意它实际上是在进行 Agentic RL,…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

1秒烧10刀! 小米正在直播模型训练! 首先注意同时在训练2个模型, mimo-v2.6-pro 和 flash, 这个无疑是进入后训练阶段了, 写着RL嘛. 然后注意它实际上是在进行 Agentic RL, 因为我们能看到衡量训练效果的测试集是DeepSWE, 不过从训练时间和得分来看, 还处于训练早期阶段. 目前mimo-v2.6-pro 的 DeepSWE得分是62, 而现在的头部模型比如DeepSeek-v4.1-flash 是 74.2. 所以训练处于相对早期. 另外注意看这个烧钱速度, 每秒至少10刀. 那么我们能否推算一下训练的用卡规模和预计的训练结束时间呢? 来. 首先注意这个数字(timing_s/outer_gen), 阶段10花费了58分钟, 这个阶段生成了多少token呢? (tokens · step 10) 22亿. 简单计算就能得出, 每秒要生成63万token. 而阶段10平均上下文长度是多少(ctx_total_length/mean)? 答案是89K. 这么长的上下文通常吞吐会下降很多, 我们按照H100来估算, 单卡吞吐在100-150tps 左右的话, 那么大概就需要4000-5000张H100. 注意这还只是训练pro用来解码的. 还有flash模型的, 算下来大概是2000-2500 张H100. 预估总卡量应该是

Posted by karminski-牙医 (39.4k followers) 2 days ago · 141 likes · 26.2k views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.3k posts from 4.7k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 21:10 UTC. Full method.