Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting.
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (
Posted by Mercor (24.3k followers) 1 h ago · 34 likes · 3.7k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- holy mog. GPT-6 Sol pricing is insane. — @Ananth7e
- GPT 6 family models — @HarshithLucky3
- It's double-release day! — @daniel_mac8
- 🚨 GPT-6 Sol is HALF the price of GPT-5.6 Sol. — @DanDr1s
- GPT 6 Sol pricing — @HarshithLucky3
- And it's live in the Codex CLI too! Now we wait for Tibo to work his magic 🔥 — @dedene
- GPT 6 Sol and Luna are out — @HarshithLucky3
- GPT-6 Sol and Luna are dirt cheap — @scaling01
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.6k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 18:12 UTC. Full method.