Claude Opus 4.8 model release
Anthropic's Claude Opus 4.8 was released yesterday. We tested it on both APEX-Agents and APEX-SWE ahead of the launch. Here's how it performed.
APEX-SWE BENCHMARK
New APEX-SWE leader
Benchmark update
Opus 4.8 (High) takes the #1 spot at 45.3% Pass@1, nearly 4 points ahead of GPT-5.3 Codex (41.5%).

Integration vs. Observability
APEX-SWE tests two categories: Integration (building and connecting systems) and Observability (diagnosing and debugging).
Opus 4.8 leads Observability by 10 percentage points at 43.3%. On Integration, it places 4th at 47.3%.
It wins overall by pairing a dominant Observability score with competitive Integration performance.

APEX-AGENTs BENCHMARK
2nd on APEX-Agents
Benchmark update
Opus 4.8 scores 42.5% Pass@1, behind Gemini 3.5 Flash (49.6%) and ahead of GPT 5.5 (38.4%).

Progress over time
Opus 4.8 (Max) is an 8.6 percentage point improvement over Opus 4.7 (Max).
In ~6 months, Opus models have improved from 18.3% to 42.5% on APEX-Agents.

Efficiency
On APEX-Agents, Opus 4.8 uses 34% fewer total tokens than GPT 5.5. It spends 36% less on context tokens and 40% more on reasoning and answer generation.
It reads less and thinks more, and still comes out ahead on efficiency.

APEX is our set of benchmarks that reflect a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work.