APEX-Agents: Management Consultant Created by experts from McKinsey, BCG, Deloitte, Accenture, EY
Show sample task for Gemini 4 Argon Gemini 4 ArgonHigh
88.4%
Show sample task for Opus 5.5 Opus 5.5Max
2 efforts80.0%
Show sample task for Opus 5.5 Show sample task for Opus 5.5 Show sample task for Sonnet 5.5 Sonnet 5.5Max
2 efforts79.7%
Show sample task for Sonnet 5.5 Show sample task for Sonnet 5.5 Show sample task for Gemini 3.8 Flash Gemini 3.8 FlashHigh
74.7%
Show sample task for Fable 5.1 Fable 5.1Max
2 efforts71.6%
Show sample task for Fable 5.1 Show sample task for Fable 5.1 Show sample task for Gemini 3.7 Flash Gemini 3.7 FlashHigh
68.8%
Show sample task for Grok 4.6 Grok 4.6xHigh
68.1%
Show sample task for GPT-6 Astra GPT-6 AstraMax
64.7%
Show sample task for Grok 4.7 Grok 4.7xHigh
63.1%
Show sample task for Opus 5 Opus 5Max
60.9%
Show sample task for Qwen3.8-Max Qwen3.8-MaxxHigh
59.7%
Show sample task for Fable 5 Fable 5Max
59.1%
Show sample task for GPT-6.1 Sol GPT-6.1 SolMax
58.4%
Show sample task for Muse Spark 1.3 Muse Spark 1.3xHigh
2 efforts57.8%
Show sample task for Muse Spark 1.3 Show sample task for Muse Spark 1.3 Show sample task for DeepSeek-V4-Flash DeepSeek-V4-FlashMax
56.7%
Show sample task for GPT-5.6 Terra GPT-5.6 TerraMax
56.2%
Show sample task for MiMo V2.6 Pro MiMo V2.6 ProAuto
56.2%
Show sample task for Grok 4.5 Grok 4.5High
55.9%
Show sample task for GPT-5.5 GPT-5.5xHigh
52.8%
Show sample task for MiMo V2.6 Flash RL MiMo V2.6 Flash RLAuto
52.4%
Show sample task for GPT-6 Sol GPT-6 SolMax
51.6%
Show sample task for Haiku 5.5 Haiku 5.5Max
50.9%
Show sample task for Sonnet 5 Sonnet 5Max
47.8%
Show sample task for GPT-5.4 GPT-5.4xHigh
44.8%
Show sample task for GLM-5.3 GLM-5.3Max
43.1%
Show sample task for GPT-5.6 Sol GPT-5.6 SolMax • Pro
42.9%
Show sample task for Opus 4.6 Opus 4.6Max
42.5%
Show sample task for GLM-5.3-Flash GLM-5.3-FlashMax
41.9%
Show sample task for Kimi K3 Kimi K3Max
40.3%
Show sample task for Qwen 3.8 27B Qwen 3.8 27BxHigh
40.0%
Show sample task for Opus 4.7 Opus 4.7Max
39.1%
Show sample task for GPT-6 Luna GPT-6 LunaMax
38.5%
Show sample task for Gemini 3.6 Flash Gemini 3.6 FlashHigh
37.8%
Show sample task for GPT-5.6 Luna GPT-5.6 LunaMax
35.1%
Show sample task for DeepSeek-V4-Pro-0813 DeepSeek-V4-Pro-0813Max
34.5%
Show sample task for Sonnet 4.6 Sonnet 4.6High
30.3%
Show sample task for Kimi K2.7 Code Kimi K2.7 CodeHigh
28.9%
Show sample task for Gemini 3.1 Pro Gemini 3.1 ProHigh
28.7%
Show sample task for MiniMax-M3 MiniMax-M3High
28.5%
Show sample task for Opus 4.8 Opus 4.8Max
28.4%
Show sample task for Inkling InklingHigh
25.5%
Show sample task for GLM-5.1 GLM-5.1
24.6%
Show sample task for GLM-5.2 GLM-5.2Max
23.8%
Show sample task for Muse Spark 1.2 Muse Spark 1.2xHigh
23.4%
Show sample task for Qwen3.5 Qwen3.5
20.0%
Show sample task for DeepSeek-V4.1-Flash DeepSeek-V4.1-FlashMax
20.0%
Show sample task for Gemini 3.5 Flash Gemini 3.5 FlashHigh
19.7%
Show sample task for Muse Spark 1.1 Muse Spark 1.1xHigh
17.9%
Show sample task for Gemini 3.5 Flash Lite Gemini 3.5 Flash LiteHigh
15.0%
Show sample task for Nemotron 3 Ultra Nemotron 3 UltraHigh
11.0%
Show sample task for DeepSeek-V3.2 DeepSeek-V3.2
8.4%
Show sample task for GPT-OSS-120B GPT-OSS-120BHigh
0.0%
0%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%
Sample task (1 out of 80) Task
Task
Environment
Environment
Trajectory
Trajectory
This sample task is provided for illustration only. The scores may not reflect an agent's overall performance on the leaderboard.
Read the attached email from the EM about conducting lifecycle analysis on the SKU data and execute on the analysis in: 4. Data Hygiene / Gaps Log (Associate 2).
Can you get back to me with the four lifecycle top-two pairs and the two most-frequent platforms/applications? Give me each lifecycle top-two as an unranked pair, so order within the pair doesn't matter, and report the final two-platform result as an unranked pair too. Supporting percentages are optional, but if you show them, round them to two decimal places.
No rubric data available for this model.