Long-horizon command-line agent execution.
Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.
Model
Score
Opus 5.5Max
Score on the original open-source benchmark: 83.5% / Score on this Mercor-extended benchmark: 47.5%
GPT-6.1 SolMax
Score on the original open-source benchmark: 85.4% / Score on this Mercor-extended benchmark: 45.8%
Grok 4.6xHigh
Score on the original open-source benchmark: 84.6% / Score on this Mercor-extended benchmark: 45.5%
Sonnet 5.5Max
Score on the original open-source benchmark: 87.3% / Score on this Mercor-extended benchmark: 45.5%
GPT-6 SolMax
Score on the original open-source benchmark: 85.4% / Score on this Mercor-extended benchmark: 42.4%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.