SITUATION EXPLAINED: Opus 5.5 beats GPT-6 Astra on most benchmarks, and costs 40% less…
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
SITUATION EXPLAINED: Opus 5.5 beats GPT-6 Astra on most benchmarks, and costs 40% less than Opus 5. • 66.4% on Terminal-Bench 4.0 against Astra's 57.9%, and 1846 on GDPval against Astra's 1542 • $4 input and $20 output, down from $5 and $25. Cache reads drop from $0.50 to $0.20, and output is 30% faster • Tested by METR and Frontier Design before release. Strongest alignment score Anthropic has posted, with 85% fewer attempts to cross containment boundaries than Opus 5 • It also often appears to suspect it's being evaluated, which Anthropic says complicates predicting how it behaves once depl
Posted by MTS (517.4k followers) 2 h ago · 33 likes · 4.4k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- if you’re using opus 5.5, steal this prompt... — @Av1dlive
- Claude Opus 5.5 system card: — @rohanpaul_ai
- anthropic won big today — @notjazii
- One launch deal for 5 new models. — @RespanAI
- Claude Opus 5.5 on Higgsfield turned one house photo and its floor plans into a full 3D… — @higgsfield_ai
- A quick recap of what we saw today, and an answer to the question of whether there was a… — @kimmonismus
- Grok 4.7 (xHigh) by @SpaceXAI just landed at #10 in Code Arena: WebDev with 1632 pts. — @arena
- Claude Opus 5.5 vs GPT-6 Sol in 3D game development. — @higgsfield_ai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.1k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 20:50 UTC. Full method.