OpenAI disputes ARC-AGI-3 benchmark setup after GPT-5.6 Sol underperforms official test

OpenAI's GPT-5.6 Sol achieved 38.3 percent on ARC-AGI-3 using proprietary API features, but scored 7.8 percent under the official test environment, undercutting Anthropic's Opus 5 record. The discrepancy raises questions about benchmark neutrality and whether the ARC Prize's testing infrastructure may rely on outdated API implementations that disadvantage certain providers. This dispute signals growing tension over how frontier models are evaluated and compared, with implications for how the industry validates capability claims.
Modelwire context
Skeptical readOpenAI's headline claim collapses under scrutiny: the 38.3% result required undisclosed API settings unavailable to the official ARC Prize evaluators. The 7.8% score under standard conditions actually loses to Anthropic. This isn't a capability win; it's a dispute over testing conditions being repackaged as a product announcement.
This is largely disconnected from recent activity in the space. We have no prior Modelwire coverage of ARC-AGI benchmark disputes or the ARC Prize testing infrastructure. What this belongs to is the broader pattern of vendors claiming benchmark superiority through methodological loopholes rather than genuine capability advances. This fits a category we should be tracking: the gap between how models perform under controlled conditions versus how vendors present them to the market.
If OpenAI publishes the exact API parameters and settings used for the 38.3% result and Anthropic or an independent lab reproduces that score, the claim gains credibility. If OpenAI declines to disclose the settings or if reproduction attempts fail, this was a marketing maneuver. The ARC Prize organizers should clarify within 30 days whether they plan to update their official test harness to include these features.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · GPT-5.6 Sol · Anthropic · Opus 5 · ARC-AGI-3 · ARC Prize
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.