Oct 1, 2026Research

Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants

Aden Barton
Aden BartonResearcher

Benchmark scores have soared over the last few years, but AI adoption seems to be contributing only slightly, if at all, to higher productivity growth in U.S. economic statistics. It's a trillion-dollar paradox.

That tension is hard to resolve without understanding how benchmarks relate to real work, and how humans themselves would do that work. To help bridge that gap, we hired 12 junior accountants to complete simplified versions of tasks from our recent APEX-Accounting benchmark.

We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.


Dot plot comparing accountants without AI and Claude Opus 5 on four accounting tasks. Accountant accuracy ranged from 0% to about 90%, and most attempts took 30 to 180 minutes. Claude scored 100% on all 20 attempts and finished each one in under 10 minutes.

More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.


Scatter plot of AI model accuracy on the study tasks by release date, from mid-2024 to late 2026. A horizontal line shows the average accountant score of 37%. GPT-4o scored near 0%. OpenAI o3 passed the accountant average in spring 2025, and GPT-5 reached about 69%. Opus 5 and other recent frontier models now score at or near 100%. Some budget and open-weight models, such as Qwen3.5-122B, still score below the accountant line.

These results do not mean accountants are replaceable, since our tasks ended up testing what AI is best at – e.g., detail-oriented instruction following – rather than the whole job of an accountant, which involves less structured tasks like client communication.

But these data do suggest substantial changes, and productivity gains, in accounting in the coming years. That would be true even if model progress froze today, and even more so if the trend depicted above continued.

These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.

Below is a brief study summary and the full paper can be found here.

What we did

Participants averaged about five and a half years of accounting experience, and all were licensed CPAs. The accountants completed four realistic month-end close scenarios, each requiring them to dig through a company's working files to find the right numbers, do the math, and deliver a table of results.

We designed the study to measure AI augmentation, but because models alone scored perfectly, there was no room to measure uplift. We focus instead on unassisted humans vs. AI.

What we found

As shown above, participants performed substantially worse than frontier models. Models are also less expensive: comparing cost per task criterion, frontier models are more than an order of magnitude cheaper than humans.


Bar chart on a log scale. Claude Opus 5 cost $0.21 per rubric criterion met. Accountants without AI cost $10.35, which is 49 times more, based on the US median accountant wage.

What do these results mean

These results don’t imply that accountants are fully replaceable, but they are informative about the future of work.

For starters, our tasks tested only a set of accounting skills that happen to be the same skills models are best at: being detail-oriented, searching through files, and following instructions very closely.

Because the tasks were originally designed to stump models, they evolved to include hard-to-spot (but realistic) accounting requirements that compounded on one another, so a single overlooked number could produce a very low score.

Being able to navigate such a complicated file set is an important skill in accounting (and knowledge work more generally), but it is not the whole job. Given this difficulty, our task authors expected average scores to be quite low: 30% for a junior and 55% for a mid-level accountant.

Further, the setting stripped out much of the job: no coworkers to ask for help, no accumulated job context. That made the tasks harder. It also meant we did not measure the skills likely to matter more as models take on the structured, detail-oriented parts of accounting: client communication, asking the right questions, building tacit context.

As one task author put it, "the traps [hidden in the task] are realistic, and the conditions and the scoring are what drove the results."

What this teaches us about benchmarking

Benchmarks aim to mirror real work, but tasks get made harder to induce model failures, like exam questions built to challenge top students. As models improve, some tasks now take many experts working together for tens of hours to build, so one person couldn't complete them without similar time and help. The accounting tasks used here are a case in point.


Two Venn diagrams of "tasks humans can perform" and "tasks AI can perform." In the traditional benchmark, all of the tasks sit where the two circles overlap. Over time, the benchmark tasks move toward the AI side, so many of them now fall outside what humans can do.

This evolution towards work that only AI can fully do makes sense, since that is likely some of the most economically valuable output. We don’t ask humans to run 60 miles an hour, yet transporting people at that speed is exactly how cars boost productivity.

That shifts the goal of frontier benchmarking from measuring how well AI does existing human work to measuring tasks that were formerly impossible or prohibitively costly. Identifying those new tasks, of course, is its own challenge.

Download the full paper here. Interested in working on similar problems or have feedback for us? Get in touch with us here.

We thank Whitney Zhang for her help in designing the study and for giving feedback on the write-up drafts.

Additionally, we thank the accounting experts who took part in the study as well as those who helped run and create the tasks and study infrastructure.