AI Radar
Support
LiveUpdated 2026-09-22 03:15 UTC

Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big…

Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big model smell"/general…Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big model smell"/general…Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big model smell"/general…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Private benchmarks useful b/c much easier to hillclimb on public ones than to have "big model smell"/general intelligence. On a benchmark of mine, new Xiaomi v2.6 Pro *literally* the worst model I've tested in two years, and the new Grok is like Luna level. Compare to top ones...

Posted by Kevin A. Bryan (23.9k followers) 1 h ago · 3 likes · 552 views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.9k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 03:15 UTC. Full method.