Modelwire
Subscribe

Willison benchmarks GPT-6 Astra reasoning levels against GPT-5.6 variants

Illustration accompanying: The Pelican comparison grid for Astra is pretty interesting

Simon Willison's hands-on comparison of GPT-6 Astra against GPT-5.6 variants reveals how OpenAI's newest reasoning model scales across inference levels. By rendering identical prompts (pelican SVGs) at low through maximum reasoning settings and measuring token consumption, Willison surfaces a practical benchmark for understanding Astra's computational tradeoffs versus prior generations. The grid exposes how reasoning depth affects both output quality and cost, offering developers concrete data for model selection in production workloads. This type of empirical, user-driven evaluation fills a gap between marketing claims and real-world performance characteristics.

Modelwire context

Analyst take

Willison's pelican grid is doing something the official Astra documentation doesn't: it puts the GPT-5.6 variants (Sol, Terra, Luna) in direct cost-per-token relief against GPT-6 Astra, giving developers a concrete basis for deciding whether Astra's reasoning depth justifies its inference overhead versus staying on a prior-generation model.

This lands in a specific context. Astra's launch was covered here through two lenses: its cybersecurity capability designation (the OpenAI Preparedness Framework story from September 1) and its staged rollout to vetted partners before public access. Neither of those pieces addressed production cost tradeoffs, which is exactly what Willison fills in. The Claude Fable 5.1 pelican story from the same day is also worth noting: Willison used the same SVG prompt format to evaluate Anthropic's model, meaning this grid is effectively one data point in a cross-vendor empirical series he's running informally. That context matters for interpreting the methodology.

Watch whether OpenAI publishes official token consumption figures for Astra's reasoning levels within the next 30 days. If they do, it will either validate or complicate Willison's grid as a reference point for developer cost modeling.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOpenAI · GPT-6 Astra · GPT-5.6 Sol · GPT-5.6 Terra · GPT-5.6 Luna · Simon Willison

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as The Pelican comparison grid for Astra is pretty interesting”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Willison benchmarks GPT-6 Astra reasoning levels against GPT-5.6 variants · Modelwire