Modelwire
Subscribe

Computer-use agent benchmarks hide massive variance in real performance

Illustration accompanying: Teach it to stop, not just to click

A new study challenges how the AI community reports agentic computer-use agent performance, revealing that single-run results mask substantial upstream variance. Researchers decomposed success-rate variance across multiple environments and found that data-draw effects and run-to-run nondeterminism dominate reported metrics, with bimodal failure distributions on harder tasks creating roughly 30% variance in outcomes from identical training. This work exposes a measurement credibility gap in agent benchmarking and suggests the field's published numbers systematically overstate reproducibility and generalization, forcing a reckoning with how RL-based agent capabilities are validated and compared.

Modelwire context

Explainer

The core provocation here is not just that benchmarks are noisy, but that the noise is structured: harder tasks produce bimodal outcomes, meaning a single reported number could represent either a near-ceiling or near-floor result depending on which random seed happened to run. That is a different and more serious problem than ordinary statistical variance.

This connects directly to the 'Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning' paper covered the same day, which derived finite-sample bounds precisely because asymptotic guarantees fail at practical scales. That paper was solving for how many interactions an agent needs before you can trust its policy identification; this paper is pointing out that the field is not even running enough trials to know whether it has reached that threshold. Together they frame a coherent problem: RL agent evaluation lacks both the theoretical scaffolding and the empirical discipline to produce numbers worth comparing. The 35B computer-use agent context also matters because commercial teams are citing single-run results to justify deployment decisions, not just academic rankings.

Watch whether any of the major agentic benchmarks (OSWorld, WebArena, or their successors) adopt multi-run reporting requirements within the next two conference cycles. If they do not, this paper's critique will remain a footnote rather than a forcing function.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

Mentions35B computer-use agent · verifier-guided repair · agentic RL

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as Teach it to stop, not just to click”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Computer-use agent benchmarks hide massive variance in real performance · Modelwire