AI Radar
Support
LiveUpdated 2026-09-22 18:12 UTC

Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting.

Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1)…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (

Posted by Mercor (24.3k followers) 1 h ago · 34 likes · 3.7k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.6k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 18:12 UTC. Full method.