Modelwire
Subscribe

AI surpasses CPAs on routine tasks but fails full accounting workflows unsupervised

Illustration accompanying: AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision

AI models have crossed a critical threshold in professional services: they now match or exceed licensed CPAs on routine accounting work in both speed and accuracy, a reversal from just 18 months prior. However, the Mercor study reveals a hard ceiling: no current model can independently complete complex end-to-end accounting tasks without human validation. This gap between narrow task mastery and holistic domain competence signals where AI augmentation will reshape professional workflows, not replace them. For accounting firms and enterprises, the strategic implication is clear: AI handles the repetitive, high-volume work while humans retain control over judgment calls and final sign-off. The finding also underscores a persistent pattern across knowledge work: benchmark wins don't yet translate to unsupervised autonomy.

Modelwire context

Analyst take

The real news isn't that AI matches CPAs on speed and accuracy (that's the headline). It's that this capability ceiling on end-to-end work creates a durable two-tier labor model where AI handles volume and humans handle liability, which means accounting firms can now plan staffing around this split rather than waiting for full autonomy.

This fits directly with the pattern from the Decoder's IT leader survey last month: organizations are deploying AI widely but seeing modest strategic impact. The accounting study shows why. AI excels at narrow task execution (matching licensed performance on routine work) but fails at holistic domain work (closing the books). That's exactly the gap identified in the model development analysis from late September, where agents generated half the proposals but humans retained 85% of decision authority. The bottleneck isn't capability on isolated tasks; it's judgment and integration. For accounting firms, this means the productivity gain comes from offloading repetitive execution, which aligns with OpenAI's framing about automating unglamorous work rather than chasing breakthrough moments.

If Mercor publishes follow-up data showing which specific judgment tasks (revenue recognition, consolidation logic, audit trail decisions) remain unsolved across all tested models, that confirms this is a durable architectural limit rather than a training data problem. If accounting firms begin hiring fewer junior associates in Q1 2027 while maintaining or growing partner headcount, that signals they're operationalizing this two-tier model.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsMercor · APEX Benchmark

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Related

Former accountant launches Tabby to automate real-time bookkeeping workflows

Autonomous AI agents reshape workplace collaboration faster than institutions adapt

WIRED - AI·

AI agents propose half of model development ideas, humans decide almost all outcomes

The Decoder·
AI surpasses CPAs on routine tasks but fails full accounting workflows unsupervised · Modelwire