New benchmark exposes how badly AI struggles with real knowledge work
Source published ·Modelwire updated
Original coverage: The Decoder ↗·How Modelwire adds context

The development
A newly released benchmark reveals a critical gap between AI model capabilities and real-world knowledge work demands, with even leading systems solving only 3 percent of realistic tasks. This finding challenges the narrative of rapid AI progress and signals that current architectures may require fundamental rethinking to handle complex, multi-step professional workflows. For enterprises betting on near-term AI productivity gains, the result underscores the distance between lab performance and deployment readiness, reshaping expectations around timeline and feasibility of knowledge-work automation.
Modelwire’s AI-generated summary of coverage from The Decoder.
Modelwire analysis
ExplainerOur AI-generated reading of the wider context and the next developments to watch.
The 3 percent figure matters less as a headline shock and more as a methodological signal: most existing benchmarks measure isolated, well-scoped tasks, and this one appears to test chained, context-dependent workflows where errors compound across steps rather than reset between questions.
This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a longer-running conversation in the AI evaluation space about whether current benchmarks actually predict deployment value. That conversation has been building across academic venues and enterprise AI teams for roughly two years, driven by repeated cases where models that score well on standardized tests fail on the messier, ambiguous inputs that real professional work generates. The benchmark described here appears to be a direct response to that criticism, attempting to operationalize 'realistic' in a way that prior evals have avoided.
Watch whether the benchmark's authors release a public leaderboard with third-party model submissions within the next 60 days. If major labs engage with it formally rather than ignoring it, that signals the methodology has enough credibility to influence internal roadmaps.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsThe Decoder
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.