New models launch almost weekly, and most rankings still judge them on chat quality or trivia-style questions.
This guide takes a different angle: it ranks the leading models on how well they handle real professional work.
It draws on Mercor's APEX AI Productivity Index: a set of benchmarks built and graded by practicing domain experts, including lawyers, investment bankers, consultants, physicians, and software engineers, who write real professional tasks and the rubric used to grade every model and agent. Rather than a single score, it measures different kinds of work
- APEX-Agents measures whether an AI agent can complete real professional tasks that span multiple steps, tools, and applications.
- APEX-SWE measures whether a model can resolve real software-engineering issues, from integration to production debugging.
Best AI models: The quick verdict
Based on APEX benchmarks results, the models that lead where the work is hardest, autonomous agents and software engineering, are Opus 5.5 and Opus 5. The table below maps the best model to each need.
| Best for | Model | Why |
|---|---|---|
| Best AI agent | Opus 5.5 (Anthropic) | Leads APEX-Agents at 73.5 |
| Best for coding | Opus 5 (Anthropic) | Leads APEX-SWE at 63.7 |
| Best open-weight | Kimi K3 (Moonshot) | 48.0 on APEX-SWE, 50.6 on APEX-Agents |
| Best value | Sonnet 5 (Anthropic) | 46.4 on APEX-SWE at roughly $2 / $20 per 1M tokens |
Every figure below uses the Pass@1 metric (first-attempt success) and, for each model, the harness that yields its highest score. APEX-Agents scores use the Loop harness; APEX-SWE scores use the Terminus-2 harness.
Best AI agent models across corporate law, investment banking, and management consulting
APEX-Agents measures whether an AI agent can complete real professional tasks that span multiple steps and applications, across corporate lawyer, investment banking analyst, and management consultant work. These are the top 3 on the default leaderboard.
| Rank | Agent | Pass@1 (Loop) | Who it is for |
|---|---|---|---|
| 1 | Opus 5.5 (Anthropic) | 73.5 | Autonomous agents on the highest-stakes professional work |
| 2 | Fable 5.1 (Anthropic) | 68.6 | Flagship agentic strength for document-heavy professional work |
| 3 | Gemini 3.7 Flash (Google) | 67.8 | The strongest non-Anthropic agent, for Google-aligned stacks |
Scores: Pass@1 on the Loop harness (loop_truncated_tools_agent).
View full APEX-Agents leaderboard →
Opus 5.5 leads clearly at the top, ahead of Fable 5.1 and Gemini 3.7 Flash.
Best AI model and agent across software engineering
APEX-SWE measures whether a model can resolve real software engineering issues, the closest of the 3 families to autonomous coding work. These are the top 3 on the default leaderboard, Pass@1.
| Rank | Model | Pass@1 | Who it is for |
|---|---|---|---|
| 1 | Opus 5 (Anthropic) | 63.7 | Flagship coding tools and the hardest engineering issues |
| 2 | Fable 5.1 (Anthropic) | 63.6 | Production coding and debugging at flagship strength |
| 3 | Grok 4.6 (xAI) | 56.4 | The strongest non-Anthropic coder, for xAI-aligned stacks |
Scores: Pass@1 on the Terminus-2 harness, the highest-scoring harness for these models.
View full APEX-SWE leaderboard →
Opus 5 and Fable 5.1 are near-tied at the top (0.1 percentage points apart), with Grok 4.6 the strongest non-Anthropic model at 56.4%.
How should you read these rankings based on your workload?
The right model depends on which one matches your workload.
- If you are deploying autonomous agents on professional work, choose based on APEX-Agents. Opus 5.5 leads there, followed by Fable 5.1 and Gemini 3.7 Flash.
- If you are building coding tools or resolving engineering issues, choose based on APEX-SWE. Opus 5 and Fable 5.1 lead there, Grok 4.6 is the strongest alternative.
- Cost or self-hosting matters most: Kimi K3 for open weights, Sonnet 5 for a low-cost hosted model with strong coding.
Notable new and rising models
The frontier is no longer just 2 US labs. 4 models stand out this cycle:
- Gemini 3.7 Flash (Google) is the strongest non-Anthropic, non-OpenAI model, scoring 67.8 on APEX-Agents, 2nd on that board.
- Grok 4.6 (xAI) is a genuine coding contender at 56.4 on APEX-SWE, and holds up on agentic work at 65.3.
- Kimi K3 (Moonshot) is the best open-weight model, at 48.0 on APEX-SWE and 50.6 on APEX-Agents, competitive with proprietary flagships on coding.
- GPT-6 Astra (OpenAI) is a strong all-rounder at 64.7 on APEX-Agents and 50.0 on APEX-SWE.
Get the full APEX benchmark data for your workflows
The public leaderboards show the aggregate scores. The full datasets are what enterprise technology teams and AI labs use to make deployment decisions, including per-domain breakdowns, harness comparisons, rubric details, and agent trajectories.
For teams that need more:
Evaluate AI model and agent performance on your own workflows with custom evaluations graded by Mercor's network of domain experts.
Get in touch →Frequently Asked Questions
What is the best AI model right now?+−
On individual leaderboards the newest models lead: Opus 5.5 tops the agentic board and Opus 5 tops software engineering.
What is the best AI model for agentic tasks?+−
On APEX-Agents, which measures agents completing multi-step professional tasks with tools, Opus 5.5 leads at 73.5, ahead of Fable 5.1 (68.6) and Gemini 3.7 Flash (67.8).
What is the best AI model for coding and software engineering?+−
On APEX-SWE scores, Opus 5 leads at 63.7, with Fable 5.1 close behind at 63.6 and Grok 4.6 the strongest non-Anthropic model at 56.4.
How often are the APEX leaderboards updated?+−
The leaderboards are updated as new frontier models are evaluated, and rankings shift with each major release. Verify scores at the time of your decision on the live APEX leaderboards.
Which scores does this article use?+−
Every figure uses the Pass@1 metric (whether a model succeeds on its first attempt) and, for each model, the harness that gives its highest score. APEX-Agents uses the Loop harness (loop_truncated_tools_agent); APEX-SWE uses the Terminus-2 harness. Per-domain scores, such as the corporate lawyer or investment banking subcategories, are not used here; those are covered in the domain-specific articles.
