Claude Opus 5 quadruples ARC-AGI benchmark record with novel reasoning behavior
Source published ·Modelwire updated
Original coverage: The Decoder ↗·How Modelwire adds context

The development
Anthropic's Claude Opus 5 has achieved a substantial breakthrough on ARC-AGI-3, a benchmark designed to measure reasoning and general intelligence. The model scored 30.2 percent, nearly quadrupling the previous record of 7.8 percent set by GPT-5.6 Sol. Notably, Opus 5 independently formulated reflection equations, a capability the benchmark's creators had not observed in prior models, suggesting a meaningful advance in logical reasoning depth. This result signals a widening capability gap between frontier labs and raises questions about how quickly reasoning-focused architectures are progressing relative to scale-driven approaches.
Modelwire’s AI-generated summary of coverage from The Decoder.
Modelwire analysis
Analyst takeOur AI-generated reading of the wider context and the next developments to watch.
The detail worth sitting with is not the score itself but the emergent behavior: Opus 5 independently formulating reflection equations suggests the model is producing reasoning strategies its own creators did not explicitly train for, which is a different kind of result than simply outperforming on a held-out test set.
Modelwire has no prior coverage to anchor this to directly, so context has to come from the broader competitive landscape this story belongs to. ARC-AGI-3 was designed specifically to resist memorization and pattern-matching, which means benchmark scores here are harder to dismiss as eval contamination than on most leaderboards. The 7.8 percent previous record attributed to GPT-5.6 Sol frames this as a two-horse race between Anthropic and OpenAI on reasoning depth, with other labs not yet visible at this tier. That framing matters because it suggests the capability gap is not just widening but concentrating.
Watch whether OpenAI responds with a targeted ARC-AGI-3 run from a GPT-5.6 variant or successor within the next 60 days. A non-response would itself be informative about whether they believe the benchmark is a meaningful signal or a distraction from their own evaluation priorities.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsAnthropic · Claude Opus 5 · GPT-5.6 Sol · ARC-AGI-3 · The Decoder
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.