Modelwire
Subscribe

IIT Bombay and Adobe reverse-engineer LLM prompts from outputs

Illustration accompanying: Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

Researchers at IIT Bombay and Adobe have demonstrated a technique that reconstructs LLM prompts from model outputs with high fidelity, bypassing the need for model weights or architecture knowledge. The method, termed Previous-Token Prediction, works across different model families and exposes a critical vulnerability for enterprises protecting proprietary system prompts. This finding reshapes threat modeling for production LLM deployments, forcing teams to reconsider what information leaks through generated text and whether prompt secrecy remains a viable security boundary.

Modelwire context

Explainer

The critical detail the summary underplays is that this attack requires only the model's output text, meaning it works against black-box API deployments where the attacker has no special access. That eliminates the most common assumption enterprises make when deciding that system prompts are safe: that obscurity through API abstraction is sufficient protection.

This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a growing body of work on LLM inference-time vulnerabilities, sitting alongside research on membership inference, training data extraction, and jailbreak transferability. The practical implication is that the security perimeter for production LLM systems has been quietly shrinking from multiple directions, and prompt confidentiality was one of the last informal assumptions still standing.

Watch whether major API providers such as OpenAI, Anthropic, or Google respond with explicit guidance or mitigations within the next 60 days. If they stay silent, that signals the industry is treating this as an acceptable risk rather than a disclosure-worthy vulnerability.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsIIT Bombay · Adobe Research · Previous-Token Prediction · The Decoder

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The Decoder originally reported this story as Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

IIT Bombay and Adobe reverse-engineer LLM prompts from outputs · Modelwire