Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants

Benchmark scores have soared over the last few years, but AI adoption seems to be contributing only slightly, if at all, to higher productivity growth in U.S. economic statistics. It's a trillion-dollar paradox.
That tension is hard to resolve without understanding how benchmarks relate to real work, and how humans themselves would do that work. To help bridge that gap, we hired 12 junior accountants to complete simplified versions of tasks from our recent APEX-Accounting benchmark.
We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.

More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.

These results do not mean accountants are replaceable, since our tasks ended up testing what AI is best at – e.g., detail-oriented instruction following – rather than the whole job of an accountant, which involves less structured tasks like client communication.
But these data do suggest substantial changes, and productivity gains, in accounting in the coming years. That would be true even if model progress froze today, and even more so if the trend depicted above continued.
These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.
Below is a brief study summary and the full paper can be found here.
What we did
Participants averaged about five and a half years of accounting experience, and all were licensed CPAs. The accountants completed four realistic month-end close scenarios, each requiring them to dig through a company's working files to find the right numbers, do the math, and deliver a table of results.
We designed the study to measure AI augmentation, but because models alone scored perfectly, there was no room to measure uplift. We focus instead on unassisted humans vs. AI.
What we found
As shown above, participants performed substantially worse than frontier models. Models are also less expensive: comparing cost per task criterion, frontier models are more than an order of magnitude cheaper than humans.

What do these results mean
These results don’t imply that accountants are fully replaceable, but they are informative about the future of work.
For starters, our tasks tested only a set of accounting skills that happen to be the same skills models are best at: being detail-oriented, searching through files, and following instructions very closely.
Because the tasks were originally designed to stump models, they evolved to include hard-to-spot (but realistic) accounting requirements that compounded on one another, so a single overlooked number could produce a very low score.
Being able to navigate such a complicated file set is an important skill in accounting (and knowledge work more generally), but it is not the whole job. Given this difficulty, our task authors expected average scores to be quite low: 30% for a junior and 55% for a mid-level accountant.
Further, the setting stripped out much of the job: no coworkers to ask for help, no accumulated job context. That made the tasks harder. It also meant we did not measure the skills likely to matter more as models take on the structured, detail-oriented parts of accounting: client communication, asking the right questions, building tacit context.
As one task author put it, "the traps [hidden in the task] are realistic, and the conditions and the scoring are what drove the results."
What this teaches us about benchmarking
Benchmarks aim to mirror real work, but tasks get made harder to induce model failures, like exam questions built to challenge top students. As models improve, some tasks now take many experts working together for tens of hours to build, so one person couldn't complete them without similar time and help. The accounting tasks used here are a case in point.

This evolution towards work that only AI can fully do makes sense, since that is likely some of the most economically valuable output. We don’t ask humans to run 60 miles an hour, yet transporting people at that speed is exactly how cars boost productivity.
That shifts the goal of frontier benchmarking from measuring how well AI does existing human work to measuring tasks that were formerly impossible or prohibitively costly. Identifying those new tasks, of course, is its own challenge.
Download the full paper here. Interested in working on similar problems or have feedback for us? Get in touch with us here.
We thank Whitney Zhang for her help in designing the study and for giving feedback on the write-up drafts.
Additionally, we thank the accounting experts who took part in the study as well as those who helped run and create the tasks and study infrastructure.