Claude Opus 5 model release
Anthropic's Claude Opus 5 debuts on the APEX leaderboards and it's the new leader on APEX-Agents at 43.5% Pass@1. It's also second on APEX-SWE at 54.7%, where it ranks first for the Integration domain.
APEX-Agents & APEX-SWE
Opus 5 ranks #1 on APEX-Agents and #2 on APEX-SWE
Top model on APEX-Agents
Opus 5 is the new top score on APEX-Agents with 43.5% Pass@1 and 60.6% mean criteria passed, which is narrowly ahead of Fable 5 at 43.3% and Muse Spark 1.1 at 41.9%. Against the prior generation, that's a clear step up from Opus 4.8's 39.3% Pass@1 and 56.2% mean.

APEX-Agents domains
The model performs well across all APEX-Agents domains, nearly matching or exceeding Fable 5 in law, consulting, and investment banking tasks.
Here's how Opus 5 compares to Fable 5:
- #1 in Law: 39.5% vs 39.1%
- #2 in Management consulting: 45.8% vs. 44.8%
- #2 in Investment banking: 45.2% vs. 46.1%

Token / Cost analysis
Opus 5 uses 45% more tokens per trajectory than Fable 5 (2.07M vs 1.43M), but Fable's per-token price is roughly double. After pricing, Opus 5 comes out about 28% cheaper per trajectory, while scoring slightly higher. Fewer tokens don't win if each one costs twice as much.
New leader on APEX-SWE Integration
Opus 5 is now the #1 ranked model for integration for APEX-SWE, which measures end-to-end system construction across heterogeneous services. Overall, Opus 5 debuts at #2 on APEX-SWE with 54.67% Pass@1. This is up 7.4 points from Opus 4.8, just behind the Fable 5 re-release at 54.8%, and ahead of Grok 4.5 at 51.2%.

APEX-SWE gains attributed to Integration
The improvements Opus 5 makes over Opus 4.8 aren't evenly distributed across the two APEX-SWE domains.
- Integration: 65.33% Pass@1 (+18.0 points)
- Observability: 44.00% Pass@1 (+0.7 points)
Opus 5 is markedly more reliable on service-heavy integration tasks like environment setup, authentication, service-URL normalization, and writing idempotent scripts consistently across runs.
