Claude Opus 4.8 tops APEX-SWE, places 2nd on APEX-Agents

Claude Opus 4.8 model release

Anthropic's Claude Opus 4.8 was released yesterday. We tested it on both APEX-Agents and APEX-SWE ahead of the launch. Here's how it performed.

APEX-SWE BENCHMARK

New APEX-SWE leader

Benchmark update

Opus 4.8 (High) takes the #1 spot at 45.3% Pass@1, nearly 4 points ahead of GPT-5.3 Codex (41.5%).

Claude Opus 4.8 at the top of the APEX-SWE leaderboard

Integration vs. Observability

APEX-SWE tests two categories: Integration (building and connecting systems) and Observability (diagnosing and debugging).

Opus 4.8 leads Observability by 10 percentage points at 43.3%. On Integration, it places 4th at 47.3%.

It wins overall by pairing a dominant Observability score with competitive Integration performance.

Claude Opus 4.8 APEX-SWE Integration and Observability scores

APEX-AGENTs BENCHMARK

2nd on APEX-Agents

Benchmark update

Opus 4.8 scores 42.5% Pass@1, behind Gemini 3.5 Flash (49.6%) and ahead of GPT 5.5 (38.4%).

Claude Opus 4.8 on the APEX-Agents leaderboard

Progress over time

Opus 4.8 (Max) is an 8.6 percentage point improvement over Opus 4.7 (Max).

In ~6 months, Opus models have improved from 18.3% to 42.5% on APEX-Agents.

Opus model progress on APEX-Agents over time

Efficiency

On APEX-Agents, Opus 4.8 uses 34% fewer total tokens than GPT 5.5. It spends 36% less on context tokens and 40% more on reasoning and answer generation.

It reads less and thinks more, and still comes out ahead on efficiency.

Claude Opus 4.8 token efficiency on APEX-Agents

APEX is our set of benchmarks that reflect a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work.