
OpenAI disputes ARC-AGI-3 benchmark setup after GPT-5.6 Sol underperforms official test
OpenAI's GPT-5.6 Sol achieved 38.3 percent on ARC-AGI-3 using proprietary API features, but scored 7.8 percent under the official test environment, undercutting Anthropic's Opus 5 record. The discrepancy raises questions about benchmark neutrality and whether the ARC Prize's testing infrastructure may rely on outdated API implementations that disadvantage certain providers. This dispute signals growing tension over how frontier models are evaluated and compared, with implications for how the industry validates capability claims.68















