Same-day Official A, thinking off: Claude Opus 5.5 hits 151/157 (96.2%). GPT-6 Luna Pro…


This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Same-day Official A, thinking off: Claude Opus 5.5 hits 151/157 (96.2%). GPT-6 Luna Pro lands 101/157 (64.3%). Same harness. Fifty-test gap. Luna is cheaper, faster, cleaner on coding syntax — then math and boxed reasoning fall over once thinking is actually off. When do you pick the fast/cheap model vs the one that finishes the harness?
Posted by Mike Gannotti (53.1k followers) 1 h ago · 4 likes · 210 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- I genuinely can’t believe I’m reposting an ASS benchmark, but this is too ridiculous not… — @TokenGremlin
- Alexandr Wang says $META is “pretty soon” dropping the most capable model it has ever… — @StockSavvyShay
- opus 5.5 medium — @tenobrus
- Elon Admitted Something He Never Said Before. — @0x0SojalSec
- Public Summary of Training Content for Grok 4.7 — @techdevnotes
- This trip to a black hole runs in a single HTML file. — @higgsfield_ai
- Claude Opus 5.5 was asked to show the inside of its mind — @MaaSonder
- Claude Opus 5.5 + Higgsfield built a 3D samurai game in one prompt. — @higgsfield_ai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.5k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 23:37 UTC. Full method.