Modelwire
Subscribe

GPT-6 Astra beats humans on reasoning, accelerates Chollet's AGI forecast

Illustration accompanying: Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

GPT-6 Astra's performance on ARC-AGI-3 marks a watershed moment in capability measurement, even as traditional benchmarks remain split on its overall standing. The model achieved human-level efficiency on a test designed to measure reasoning rather than scale, prompting François Chollet to accelerate his AGI timeline by roughly half. This divergence between benchmarks and specialized reasoning tasks signals a shift in how the field should evaluate progress: raw scores may obscure genuine breakthroughs in problem-solving efficiency that matter more to AGI forecasting than aggregate leaderboard position.

Modelwire context

Analyst take

The more consequential detail the summary underplays is that Chollet accelerating his forecast is not a casual endorsement. ARC-AGI-3 was explicitly designed to resist the scaling tricks that inflate conventional benchmark scores, so a human-efficiency result there carries a different evidentiary weight than leaderboard gains elsewhere.

This lands directly on top of Hugging Face's BenchMIRT investigation from September 1st, which argued that most benchmarks measure narrow task performance rather than genuine reasoning, and that the gap between scores and real capability is systematically misleading. ARC-AGI-3's divergence from traditional benchmarks is almost a live demonstration of BenchMIRT's thesis. It also sits in tension with the Claude Fable 5.1 coverage from The Decoder, where Anthropic's science benchmark jump was framed as a differentiator. If ARC-AGI-3 becomes the credible signal and Fable 5.1 underperforms on it, Anthropic's positioning weakens faster than its leaderboard standing suggests.

Watch whether Epoch AI publishes an independent replication of GPT-6 Astra's ARC-AGI-3 efficiency results within the next 60 days. If the numbers hold under third-party conditions, Chollet's revised timeline will be hard to dismiss; if they don't replicate, the divergence between benchmarks is noise rather than signal.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOpenAI · GPT-6 Astra · Claude Fable 5.1 · François Chollet · ARC-AGI-3 · Epoch AI

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The Decoder originally reported this story as Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

GPT-6 Astra beats humans on reasoning, accelerates Chollet's AGI forecast · Modelwire