A maintainer-built benchmark of 150 real coding tasks (Main and Extended) from 36 flagship open-source repositories. Unlike typical coding benchmarks that only check whether tests pass, it grades whether a maintainer would merge the patch: correctness, tests, scope, style, and codebase conventions.
Cognition created FrontierCode to evaluate whether coding agents can write code that is good enough to merge into a production codebase, not just code that passes tests. The tasks require agents to infer what the maintainer wants from a short brief, then deliver a patch that is correct, well tested, tightly scoped, and consistent with the codebase's conventions.
FrontierCode was built in three steps. First, maintainers of the repositories picked real issues from their own repos and wrote short, humanlike task briefs, spending more than 40 hours on each task. Second, they wrote grading criteria that combine unit tests, scope checks, and rubrics, and marked which criteria would block a merge. Third, each criterion went through adversarial testing, calibration, and multi-stage review, and a Cognition researcher checked every task by hand.
There are 150 tasks in FrontierCode, drawn from 36 open-source repositories and graded by more than 1,000 criteria. Main is the 100 hardest tasks, and Extended is the full set of 150. The tasks stay private to prevent contamination, and Cognition opens evaluation to all model creators.
FrontierCode is a coding benchmark from Cognition that tests whether AI agents can write code a maintainer would actually merge. Each task pairs a real open-source repository with a single issue. The agent works on its own in a container and submits a patch. The patch is graded on correctness, test quality, scope, style, and how well it follows the codebase's standards.
More than 20 open-source maintainers built the tasks from repositories they maintain, including Celery, Budibase, Uppy, and Mattermost. Each maintainer spent more than 40 hours per task and went through several rounds of review with Cognition researchers. The maintainers define what "mergeable" means in their own repo.
In one C++ task from the jsonschema repository, the agent must add a new LOG_WARNING() helper and use it for every warning message in the codebase. The task looks simple, but models often fail one blocker criterion. For multi-line warnings, they call LOG_WARNING() once and then write the remaining lines straight to std::cerr. The output looks the same. The code still breaks if someone changes LOG_WARNING() later, so a maintainer would not merge it.
Each task has blocker and non-blocker criteria. Blockers are the hard stops in a real code review, such as correctness, regressions, or touching files the patch should not touch. Non-blockers are quality signals such as style, type safety, and readability. A patch passes only if it clears every blocker. Its score is the weighted total of the criteria it meets. A patch that fails any blocker scores 0. Each model runs 5 times at each reasoning effort, and the leaderboard reports the score at its best effort.
It uses six types of checks. Injected unit tests check behavior. Build and lint commands check for regressions and clean code. Reverse tests run the agent's own tests on the original, broken code, where they must fail, which proves that the tests are meaningful. Adaptive tests change the reference tests to fit valid alternative solutions. Scope checks limit which files and how many lines a patch can change. An LLM reviewer grades code quality against the maintainer's written standard.
Main has the 100 hardest tasks. Extended has all 150 tasks. Cognition first reported a third subset, Diamond, with the 50 hardest tasks. FrontierCode 1.1 retired Diamond because the very low solve rates on those tasks made its results noisy.
FrontierCode scales difficulty through quality standards, not patch size. A patch that works but adds unneeded refactors, skips meaningful tests, or ignores project conventions can still fail. Even the strongest models miss on many tasks, because passing tests is only the minimum.
Yes, the way an engineer would: reading documentation, checking API references, and searching error messages. Some tasks need internet access by design. Since FrontierCode 1.1, agents get a clear "fair internet use" prompt, and a verifier flags any run that opens solution-bearing sources such as the original pull request. Flagged runs score 0.
Both come from work with Cognition, but they measure different things. APEX-SWE, built by Mercor in partnership with Cognition, tests integration work across services and debugging with production-style telemetry. FrontierCode tests whether an agent's fix to a single real open-source issue meets the maintainer's bar for merging.
Rankings can change whenever a new frontier model comes out. See the leaderboard above for current results.
No. Cognition keeps the tasks private to prevent contamination. Model creators can request an evaluation from Cognition. The FrontierCode leaderboard on Cognition's site shows a sample task you can explore.