Researchers extract hidden reasoning from frontier LLM APIs via replay attacks
Source published ·Modelwire updated
Original coverage: Simon Willison ↗·How Modelwire adds context

The development
Researchers have demonstrated a practical attack against reasoning transparency features deployed by Anthropic, OpenAI, and Google. By extracting encrypted chain-of-thought traces from API responses and replaying them into weaker model variants, attackers can decrypt and expose the internal reasoning of frontier models. This finding exposes a fundamental tension in the current approach to interpretability: making reasoning visible for safety and auditability creates a new attack surface for model extraction and jailbreaking. The vulnerability suggests that frontier labs may need to rethink how reasoning artifacts are handled across API boundaries.
Modelwire’s AI-generated summary of coverage from Simon Willison.
Modelwire analysis
ExplainerOur AI-generated reading of the wider context and the next developments to watch.
The attack doesn't require breaking encryption directly. It exploits the fact that weaker model variants, which share architectural lineage with frontier models, can serve as decryption oracles when fed the same encrypted trace artifacts, meaning the vulnerability is structural to how these labs have deployed reasoning families rather than a flaw in any single model.
Modelwire has no prior coverage to anchor this to directly, so context has to come from the broader interpretability debate. Frontier labs have spent the last two years arguing that visible reasoning is a safety property, the idea being that auditable thought processes let researchers catch misalignment before it causes harm. This finding complicates that framing considerably: the same artifact that enables oversight also encodes enough signal about the frontier model's internal behavior to assist extraction and jailbreak attempts. The tension here is not new in security research (audit logs have always been a target) but it is new at this layer of the AI stack.
Watch whether Anthropic, OpenAI, or Google quietly deprecate or scope-limit encrypted trace access in their APIs within the next 60 days. A silent policy change would confirm the labs consider this a live threat rather than a theoretical one.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsAnthropic · OpenAI · Google · Simon Willison
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as “Stealing Reasoning Traces from Proprietary LLM APIs”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.