Modelwire
Subscribe

Researchers extract hidden reasoning from frontier LLM APIs via replay attacks

Illustration accompanying: Stealing Reasoning Traces from Proprietary LLM APIs

Researchers have demonstrated a practical attack against reasoning transparency features deployed by Anthropic, OpenAI, and Google. By extracting encrypted chain-of-thought traces from API responses and replaying them into weaker model variants, attackers can decrypt and expose the internal reasoning of frontier models. This finding exposes a fundamental tension in the current approach to interpretability: making reasoning visible for safety and auditability creates a new attack surface for model extraction and jailbreaking. The vulnerability suggests that frontier labs may need to rethink how reasoning artifacts are handled across API boundaries.

Modelwire context

Explainer

The attack doesn't require breaking encryption directly. It exploits the fact that weaker model variants, which share architectural lineage with frontier models, can serve as decryption oracles when fed the same encrypted trace artifacts, meaning the vulnerability is structural to how these labs have deployed reasoning families rather than a flaw in any single model.

Modelwire has no prior coverage to anchor this to directly, so context has to come from the broader interpretability debate. Frontier labs have spent the last two years arguing that visible reasoning is a safety property, the idea being that auditable thought processes let researchers catch misalignment before it causes harm. This finding complicates that framing considerably: the same artifact that enables oversight also encodes enough signal about the frontier model's internal behavior to assist extraction and jailbreak attempts. The tension here is not new in security research (audit logs have always been a target) but it is new at this layer of the AI stack.

Watch whether Anthropic, OpenAI, or Google quietly deprecate or scope-limit encrypted trace access in their APIs within the next 60 days. A silent policy change would confirm the labs consider this a live threat rather than a theoretical one.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsAnthropic · OpenAI · Google · Simon Willison

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as Stealing Reasoning Traces from Proprietary LLM APIs”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Researchers extract hidden reasoning from frontier LLM APIs via replay attacks · Modelwire