FrontierSWE v2

Ultra-long-horizon software engineering.

20 hagent budget per task
5categories
34Public tasks
37.2%Highest score

The FrontierSWE v2 leaderboard

FrontierSWE v2 asks agents to complete projects that take expert engineers days. Examples include porting git to Zig, writing a Lean 4 kernel type checker in Pascal, optimizing FFmpeg's libswscale, and training a speech decoder from MEG signals. Every task is scored continuously against published or measured anchors, so models earn partial credit for partial progress. It remains a significant unsaturated public coding benchmark.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.