Skip to content
Modelwire
Subscribe

IIT Bombay and Adobe reverse-engineer LLM prompts from outputs

Source published ·Modelwire updated

Original coverage: The Decoder ↗·How Modelwire adds context

Illustration accompanying: Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

The development

Researchers at IIT Bombay and Adobe have demonstrated a technique that reconstructs LLM prompts from model outputs with high fidelity, bypassing the need for model weights or architecture knowledge. The method, termed Previous-Token Prediction, works across different model families and exposes a critical vulnerability for enterprises protecting proprietary system prompts. This finding reshapes threat modeling for production LLM deployments, forcing teams to reconsider what information leaks through generated text and whether prompt secrecy remains a viable security boundary.

Modelwire’s AI-generated summary of coverage from The Decoder.

Modelwire analysis

Explainer

Our AI-generated reading of the wider context and the next developments to watch.

The critical detail the summary underplays is that this attack requires only the model's output text, meaning it works against black-box API deployments where the attacker has no special access. That eliminates the most common assumption enterprises make when deciding that system prompts are safe: that obscurity through API abstraction is sufficient protection.

This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a growing body of work on LLM inference-time vulnerabilities, sitting alongside research on membership inference, training data extraction, and jailbreak transferability. The practical implication is that the security perimeter for production LLM systems has been quietly shrinking from multiple directions, and prompt confidentiality was one of the last informal assumptions still standing.

Watch whether major API providers such as OpenAI, Anthropic, or Google respond with explicit guidance or mitigations within the next 60 days. If they stay silent, that signals the industry is treating this as an acceptable risk rather than a disclosure-worthy vulnerability.

This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error

MentionsIIT Bombay · Adobe Research · Previous-Token Prediction · The Decoder

MW

How this coverage is produced

Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.

Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

IIT Bombay and Adobe reverse-engineer LLM prompts from outputs · Modelwire