Anthropic's Claude Fable 5 was released this week. We tested it on both APEX-Agents and APEX-SWE ahead of the launch. Here's how it performed.
APEX-SWE
Pass@1, at release
Fable 5 — 65.5% ± 6.2%
Opus 4.8 (High) — 45.3% ± 6.3%
GPT 5.3 Codex (High) — 41.5% ± 6.3%
Opus 4.7 (Max) — 41.3% ± 6.3%
GPT 5.5 (xHigh) — 40.8% ± 6.5%
Fable 5 by APEX-SWE domain
Observability — 69.7%
Integration — 61.3%
APEX-Agents
Score, at release
Gemini 3.5 Flash (High) — 49.6% ± 3.9%
Fable 5 (Max) — 45.0% ± 4.1%
Opus 4.8 (Max) — 42.5% ± 4.0%
GPT 5.5 (xHigh) — 38.4% ± 3.9%
GPT 5.4 (xHigh) — 36.0% ± 3.8%
Average token usage
Claude Fable 5 (Max) — 924k tokens per APEX-Agents run
By domain
Corporate Lawyer Pass@1 — Claude Fable 5 (Max) 40.9%
