Modelwire
Subscribe

OpenART exposes agent safety gaps beyond isolated task benchmarks

Researchers have built OpenART, a large-scale red teaming framework that exposes a critical gap in how AI agents are currently evaluated. Most safety benchmarks test agents on isolated tasks, but real-world deployment involves persistent environments where early actions compound into downstream consequences. OpenART's 10,000+ stateful scenarios across 50 domains force agents to navigate cumulative risk over long workflows, requiring median 97-step tool chains. This work signals that agent safety evaluation is fundamentally different from language model benchmarking, and existing threat models may be blind to emergent failure modes in multi-step, state-dependent reasoning.

Modelwire context

Analyst take

OpenART reveals that existing agent benchmarks are fundamentally misaligned with how deployed systems actually fail. The gap isn't about isolated task performance but about state-dependent failure modes that compound across long workflows, a problem that isolated benchmarks cannot surface.

This lands directly in the tension between OpenAI's Presence launch (moving agents to production) and the safety infrastructure catching up. OpenAI is reportedly building Astra for multi-day reasoning tasks, yet CompressAgent's August 2nd findings show that even modest context compression creates unpredictable reliability degradation in tool-using agents. OpenART now documents that 97-step workflows expose emergent failure modes invisible to standard evals. Together, these three stories suggest the industry is shipping production agent infrastructure faster than it can measure whether those systems are actually safe in stateful, long-horizon settings.

If OpenAI's Presence customers report cascading failures or unexpected agent behavior within the first 90 days of deployment on workflows longer than 20 steps, that confirms OpenART's warning. Conversely, if Presence ships with built-in stateful scenario testing derived from OpenART's framework, that signals the labs are already incorporating these findings into production readiness criteria.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOpenART

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Related

OpenAI develops Astra for multi-day agent reasoning tasks

The Decoder·

OpenAI's Astra tackles unsolved math via multi-agent reasoning

The Decoder·

OpenAI launches Presence to operationalize AI agents for enterprises

The Decoder·
OpenART exposes agent safety gaps beyond isolated task benchmarks · Modelwire