Coding assistants drastically overestimate task duration and self-performance

A new study reveals that AI coding assistants systematically misjudge temporal constraints and their own output quality, with Codex overestimating task duration by up to tenfold and both models inflating performance self-assessments by roughly 20 percentage points. This gap between perceived and actual performance creates meaningful governance challenges for autonomous systems operating without human oversight, particularly in production environments where miscalibrated time estimates and inflated confidence could compound into costly failures or safety incidents.
Modelwire context
ExplainerThe study isolates a distinct failure mode: these models don't just perform poorly on hard tasks, they actively misread how long tasks take and overstate their own accuracy. This is different from general capability gaps because it means a system could confidently commit to impossible deadlines or ship code it believes is correct when it isn't.
This is largely disconnected from recent activity in the space, which has focused on scaling, reasoning benchmarks, and alignment. The temporal miscalibration finding belongs to a smaller but growing body of work on model self-awareness and calibration (how well a model's confidence matches reality). It matters because autonomous deployment assumes some baseline honesty in self-assessment. If systems are systematically overconfident about both speed and quality, the governance frameworks being built around 'human-in-the-loop' oversight become fragile when humans trust the system's own time and confidence signals.
If Claude or Codex show the same temporal miscalibration pattern when tested on the same benchmark, that confirms this is a structural property of how these models process time, not a quirk of one architecture. If vendors begin publishing calibration metrics (confidence vs. accuracy curves) as standard practice in the next 12 months, that signals the industry is taking this seriously.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsClaude · Codex · The Decoder
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “AI agents have no sense of time and are not aware of it”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.