There are some new top models leading our code review benchmark.
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
There are some new top models leading our code review benchmark. 🥇 Claude Opus 5.5 — 82.3 (f1 score) 🥈 GPT-6 Sol — 80.3 🥉 GPT-6 Astra — 78.0 See full results here: https://www.macroscope.com/benchmark
Posted by Macroscope (1.2k followers) 1 h ago · 6 likes · 469 views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: Macroscope.
More AI work like this
- Gave Opus 5.5 a try with a trailer with my web game… — @decruz
- I evaluated the latest models, GPT-6 Sol, Luna, and Claude Opus 5.5 on induction (post… — @s_batzoglou
- Claude Opus 5.5 is the greatest AI model ever released — @AlexFinn
- GLM-5.3 beat Space Bunny Alpha (the new stealth model launched today) on Newton’s cradle… — @rohanpaul_ai
- The token usage chart here is really striking — @charliermarsh
- Anthropic’s guide to using Opus 5.5 has one central rule: stop asking it to think. Eddie… — @davidarngar
- Claude Opus 5.5 recreated an athlete’s long jump in 4D. — @higgsfield_ai
- Claude Opus 5.5 painted the past and future of our world in JavaScript + Higgsfield. — @higgsfield_ai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.7k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 02:07 UTC. Full method.