DAYJOB benchmark exposes wide gap in AI professional task performance
Researchers have released DAYJOB, a benchmark that measures how well AI systems handle realistic professional workflows in healthcare and finance. The 130 tasks, each requiring 13-17 hours of expert work, reveal a significant capability gap: even the strongest model (Claude Opus 5.5) succeeds on fewer than 25% of attempts, while median systems barely exceed 0.6%. The benchmark uses containerized environments and strict multi-criterion evaluation, making it a more rigorous test than typical academic benchmarks. This work signals growing focus on evaluating agentic systems against real-world complexity rather than isolated task performance, and suggests current models remain far from autonomous professional deployment.
Modelwire context
Analyst takeThe benchmark's real significance isn't the low absolute scores (which confirm what safety researchers already knew) but the specificity of failure modes in containerized, multi-step workflows. DAYJOB measures whether models can sustain coherence across 13-17 hour tasks with real environmental feedback, not whether they can solve isolated problems.
This arrives as OpenAI and others are shipping production agentic infrastructure (Computer Use, Agents API from late September) and enterprises are beginning deployment (Standard Chartered's security agents, the WIRED piece on workforce integration). DAYJOB functions as a reality check: it measures whether the infrastructure improvements announced at DevDay actually translate to professional-grade autonomy. The gap between Claude Opus 5.5's 25% success rate and the deployment velocity described in recent coverage suggests organizations are moving faster than capability justifies, creating a window where early adopters will face significant failure rates before models mature.
If Anthropic or OpenAI release updated benchmarks on DAYJOB within six months showing >40% success on the same tasks, that signals either genuine capability acceleration or benchmark saturation (contamination). If enterprise deployments of agentic systems in finance and healthcare remain confined to narrow, human-supervised workflows through 2027, that confirms DAYJOB's findings are predictive of real-world constraints rather than academic artifacts.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsClaude Opus 5.5 · DAYJOB · Anthropic
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “DAYJOB: A Benchmark for Long-Horizon Professional Work”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.