Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big…


This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big model smell"/general intelligence. On a benchmark of mine, new Xiaomi v2.6 Pro *literally* the worst model I've tested in two years, and the new Grok is like Luna level. Compare to top ones...
Posted by Kevin A. Bryan (23.9k followers) 1 h ago · 3 likes · 552 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- 如果真的是这样,我建议马斯克也做中转站,直接去小米进货,然后把 grok 路由到 MiMo,价格差了 14 倍,比你卖车赚钱多了🤣 — @dingyi
- 플루토 리서치에서 퍼왔습니다. 정말 이렇게 많이 차이가 나나요? — @joejo2038
- 感觉现在发模型只能走两个极端: — @op7418
- Congrats to Grok 4.7 and Xiaomi’s MiMo-2.6 Pro & MiMo-2.6 Flash 🎉 — @YouWareAI
- Grok 4.7 improves from 54% to over 60% on Vals Index after updating the SDK — @scaling01
- 小米正式发布并开源 MiMo-V2.6 系列,包含 Pro 和 Flash 两个原生多模态模型。 — @dotey
- New cheap highly effective model just dropped. New frontier — @Perpetualmaniac
- Grok 4.7 vs 4.6, round 2. — @WescheNex1q
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.9k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 03:15 UTC. Full method.