GPT-6 Astra outpaces rivals in dual-arm robotics reasoning

OpenAI's GPT-6 Astra demonstrates a meaningful leap in embodied AI by solving complex dual-arm robotics tasks where competitors fail entirely. On StationeryBench, the model completed 7 of 100 tasks while MolmoAct2 scored zero, signaling progress toward practical robot control beyond simulation. This capability jump matters because spatial reasoning and multi-limb coordination have been persistent bottlenecks in deploying language models to physical systems. The result suggests frontier labs are cracking the integration layer between language understanding and real-world manipulation, a prerequisite for autonomous systems at scale.
Modelwire context
Skeptical readSeven out of 100 completed tasks is the actual number being called a 'step change,' and the comparison baseline is a model that scored zero, not a prior version of GPT or a broadly accepted robotics benchmark. That framing makes the delta look larger than the absolute performance warrants.
This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a thread running through the broader embodied AI space, where labs have repeatedly announced spatial reasoning improvements tied to proprietary benchmarks that don't replicate cleanly in third-party settings. StationeryBench appears to be a new evaluation surface, which means there is no longitudinal baseline to judge whether 7 percent completion is impressive or a floor.
If an independent robotics lab reproduces the StationeryBench results on their own hardware within the next 90 days, the capability claim gains real weight. If the only published scores continue to come from OpenAI-adjacent sources, treat this as an internal milestone rather than a field-wide advance.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · GPT-6 Astra · MolmoAct2 · StationeryBench
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.