Modelwire
Subscribe

OpenAI's Sol model shows 38% on custom tests, 7.8% on standard ARC-AGI-3

Illustration accompanying: OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI's latest frontier model, GPT-5.6 Sol, reveals a widening gap between lab conditions and real-world benchmarking. The model achieves 38.3 percent on ARC-AGI-3 using OpenAI's proprietary test harness with extended reasoning, but drops to 7.8 percent in the official environment, where Anthropic's Opus 5 scores 30.2 percent without such scaffolding. This discrepancy signals a critical tension in how frontier labs measure progress: custom infrastructure can inflate apparent capability gains, while standardized benchmarks remain the only reliable cross-vendor comparison. For practitioners and investors, the gap underscores that claimed breakthroughs require scrutiny of test conditions before informing deployment decisions.

Modelwire context

Skeptical read

OpenAI is not claiming a general capability breakthrough on ARC-AGI-3. It is claiming that its custom reasoning harness can extract higher performance from the same model on the same task, a methodological advantage rather than a model advantage. The real story is that this gap (38.3% vs 7.8%) exposes how easily infrastructure choices can manufacture apparent progress.

This is largely disconnected from recent activity in the space, as we have no prior coverage of ARC-AGI-3 benchmarking or the specific Opus 5 vs GPT-5.6 Sol comparison. However, it belongs to an ongoing pattern in frontier model evaluation: the tension between what labs can demonstrate in controlled settings versus what independent benchmarks reveal. This is the first time we're covering the specific mechanics of that gap for this generation of models.

If Anthropic releases its own extended reasoning harness for Opus 5 on ARC-AGI-3 and closes the gap to within 5 percentage points of OpenAI's 38.3% score, that confirms the difference is infrastructure, not model quality. If the gap persists, it suggests OpenAI's scaffolding is genuinely extracting capabilities Opus 5 lacks. Either outcome should be public within 60 days.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOpenAI · GPT-5.6 Sol · Anthropic · Opus 5 · ARC-AGI-3

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The Decoder originally reported this story as OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

OpenAI's Sol model shows 38% on custom tests, 7.8% on standard ARC-AGI-3 · Modelwire