Researchers extract hidden reasoning from OpenAI, Anthropic, and Google APIs

Researchers uncovered a cross-platform vulnerability affecting OpenAI, Anthropic, and Google that exposes encrypted reasoning traces through their APIs, enabling extraction and transfer of internal model computations. Public session scans revealed dozens of exposed credentials and API keys, while also demonstrating a critical gap between user-facing reasoning summaries and actual model behavior. This finding exposes both immediate security risks and a deeper transparency problem: the reasoning outputs users trust may systematically misrepresent what models are computing internally, raising questions about auditability and control in production AI systems.
Modelwire context
ExplainerThe credential exposure is the attention-grabbing detail, but the more durable finding is structural: the reasoning summaries that users and enterprises rely on to audit model behavior may not correspond to what the model is actually computing, which means current approaches to AI oversight built on inspecting those summaries are working from incomplete information.
This lands at an awkward moment for OpenAI specifically. Brad Lightcap's departure, covered here on August 11th, already raised questions about organizational coherence at a company managing both frontier research and production infrastructure at scale. A cross-platform vulnerability of this kind is precisely the sort of operational failure that a COO-level function is supposed to catch before it reaches public session scans. The River AI funding story from the same day is less directly connected, though it reinforces a broader pattern: the trust deficit accumulating around incumbent labs is part of what makes a $1.1B bet on alternative agent infrastructure feel rational to investors right now.
Watch whether OpenAI, Anthropic, or Google publish a formal disclosure or patch timeline within the next 30 days. Silence or a vague 'we take security seriously' response would confirm that the transparency gap identified in the research extends to incident response as well.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · Anthropic · Google · ChatGPT · The Decoder
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “"But marinade" and leaked passwords are what researchers found in ChatGPT's hidden reasoning”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.