Long-horizon command-line agent execution.
Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.
Model
Score
DeepSeek-V4.1-FlashMax
86.5%
GPT-6 AstraxHigh
85.4%
GPT-5.6 SolxHigh
84.6%
Grok 4.6xHigh
Opus 5High
83.9%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.