AI Beats CPAs at Bookkeeping, but Still Needs Oversight

AI Beats CPAs at Bookkeeping, but Still Needs Oversight

AI models now handle structured bookkeeping work faster, more accurately and far more cheaply than trained accountants. That is the main result of a new study from Mercor. The same study also shows that these systems still cannot finish a full accounting cycle without a human checking the work.

The two findings point in different directions, and both matter. Below is what the study measured, what it left out, and why the gap between the two is the real story.

The head-to-head test

Mercor recruited 12 licensed CPAs. CPA stands for certified public accountant, the main professional accounting credential in the United States. On average, the participants had five and a half years of experience. They worked through simplified tasks taken from the APEX Accounting Benchmark, Mercor's test suite for accounting work.

The human average on these tasks was about 37 percent. Eighteen months ago, even the strongest AI models scored below that line. Today, according to Mercor, the best models solve the same set of tasks almost without error.

In that window, the machines went from trailing working professionals to clearly outperforming them on this narrow slice of the job. Mercor reports that the models were also faster and cost much less per task.

The full benchmark tells a different story

The simplified tasks are only a small part of APEX Accounting. The complete benchmark contains 160 tasks spread across 10 simulated companies. More than 40 professionals built it, and they average 11 years of experience in the field.

On this harder version, current scores look like this:

  • Claude Opus 5.5: 61.8 percent of grading criteria met
  • Fable 5.1: 61.0 percent
  • GPT-6 Astra: 57.9 percent

These numbers measure grading criteria, not finished tasks. Mercor says no model fully completed almost 60 percent of the tasks in the benchmark. A model can get most of the individual steps right and still leave a task unfinished.

That is why Mercor concludes that AI cannot yet close the books on its own. Closing the books is the period-end process where a company finalizes its accounts for a month, quarter or year. It requires every piece to be correct, not most of them.

What the study did not test

Mercor is open about the limits of its own experiment. The company says the tasks were close to ideal for AI. They reward two things models are very good at: tracking down specific details and following instructions to the letter.

Several core parts of an accountant's job were left out entirely:

  • Conversations with clients
  • Coordination with colleagues
  • Judgment built on years of context about a business

Mercor points to these missing pieces as the reason accountants are not about to be replaced. At the same time, it expects AI to bring major productivity gains across the accounting industry.

Reading the numbers carefully

The study supports two claims, and it is easy to mix them up.

The first is narrow and strong. On well-defined bookkeeping tasks, current models beat licensed professionals on speed, accuracy and cost. The jump from below 37 percent to near-perfect in 18 months is large.

The second is broad and cautious. On a more realistic benchmark, the best model meets only about 62 percent of grading criteria, and no model fully solves most tasks. The work that remains is the work that ties everything together.

Headlines that report only the first claim overstate what AI can do in an accounting department today. Headlines that report only the second miss how quickly the structured part of the job has changed.

The Bigger Picture

For teams thinking about AI in finance, this study suggests a clear split. Routine, rule-bound bookkeeping looks ready for heavy automation, with humans reviewing outputs rather than producing them. End-to-end processes such as closing the books still need a person in charge.

This fits a wider pattern. Models increasingly beat experts on clean, isolated tasks, while full workflows with many dependent steps remain unreliable. The same tension shows up across the push into enterprise AI automation, where vendors promise end-to-end handling of business processes.

It is also worth noting who ran the test. Mercor built the benchmark and published the study, so independent replication would add weight to its results.

What to watch next is the gap between the 37-percent-style tasks and the full benchmark. If the full APEX scores climb as fast as the simplified ones did, the case for human oversight will get narrower. Cost will matter too, as cheaper models such as OpenAI's GPT-6.1 Sol push the price of this work down further. For now, the evidence points to accountants supervising AI, not being replaced by it.