Mercor logo
APEX
APEX-AgentsAPEX-AccountingAPEX-SWEAPEXOff-the-shelf data
Research
Enterprise
Enterprise agentsHuman dataData monetization
ExpertsMission
Log in
APEX
Research
Enterprise
ExpertsMission
  1. APEX Newsletter
  2. Model releases

Model releases

AllModel releases
Introducing new APEX-Accounting benchmark, built with Ramp
Model releases1 mins read

Introducing new APEX-Accounting benchmark, built with Ramp

APEX-Accounting is live. No frontier model can reliably close the books.

Model releases1 mins read

Opus 5 tops the APEX-Agents leaderboard

Claude Opus 5 debuts at #1 on APEX-Agents and takes the top spot on SWE Integration.

Model releases3 mins read

Grok 4.5 and GPT-5.6 land on APEX

Grok 4.5 debuts at #2 on APEX-SWE and takes the overall Integration lead.

Model releases1 mins read

Fable 5 is back. Here's what changed.

Fable 5's re-release scores 54.8% on APEX-SWE, still 9.5 points ahead of Opus 4.8.

Model releases1 mins read

Claude Sonnet 5 earns #3 spot on APEX-SWE and makes APEX-Agents top 10

Claude Sonnet 5 ranks #3 on APEX-SWE (43.7%) and jumps into the APEX-Agents top 10.

Model releases1 mins read

Claude Fable 5 tops APEX-SWE with a 20-point lead

Anthropic's Claude Fable 5 was released this week. We tested it on APEX-Agents and APEX-SWE ahead of the launch.

Model releases1 mins read

Claude Opus 4.8 tops APEX-SWE, places 2nd on APEX-Agents

Anthropic's Claude Opus 4.8 was released yesterday. We tested it on APEX-Agents and APEX-SWE ahead of the launch.

Model releases1 mins read

Gemini 3.5 Flash tops all APEX-Agents leaderboards

Google's Gemini 3.5 Flash (High) took the top spot on the APEX-Agents leaderboard with 49.6% at Pass@1.

Experts
Find workHelp centerResourcesStories
Research
APEX BenchmarksAPEX-AgentsAPEX-AccountingAPEX-SWEOff-the-shelf data
Enterprise
Enterprise agentsHuman dataData monetization
Contact
SupportPressSales
Mercor
CareersSecurityBlogNewsroom
© 2026 Mercor
San Francisco, CA
© 2026 Mercor