EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

Researchers have constructed EnterpriseClawBench, a benchmark derived from real workplace agent sessions that exposes a significant capability gap in current enterprise AI systems. The dataset comprises 852 reproducible tasks extracted from proprietary workplace interactions, complete with fixtures, prompts, and semantic evaluation rubrics. The benchmark reveals that even state-of-the-art configurations (Codex with GPT-5.5) achieve only 66.3% success, signaling that enterprise-grade agent deployment remains substantially harder than consumer-facing LLM applications. The reusable evaluation protocol, rather than the withheld data itself, constitutes the contribution, establishing a new standard for measuring agent performance in heterogeneous, tool-heavy business environments.
Modelwire context
Analyst takeThe 66.3% ceiling matters less than what it measures: the benchmark is built from actual workplace sessions, meaning failure modes here reflect real organizational friction (heterogeneous tools, ambiguous business context) rather than the clean synthetic tasks that inflate most agent leaderboards. The withheld data is a deliberate moat, making the evaluation protocol itself the distributable asset.
This connects directly to the MAS-PromptBench work published the same day, which asks whether prompt optimization actually improves multi-agent performance in production-like settings. EnterpriseClawBench provides exactly the kind of grounded, tool-heavy evaluation surface that MAS-PromptBench lacks: if prompt tuning gains don't hold on enterprise task distributions, that finding has real procurement implications. Together, these two papers sketch a more honest picture of where agentic systems actually stand versus where benchmark-optimized demos suggest they are.
Watch whether any major enterprise software vendor (Salesforce, ServiceNow, Microsoft) formally adopts the evaluation protocol within the next two quarters. Adoption by one named platform would signal the benchmark is becoming a procurement reference point rather than staying an academic artifact.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsEnterpriseClawBench · Codex · GPT-5.5
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.