Modelwire
Subscribe

Anthropic's Opus 5 shows marked resistance to prompt injection attacks

Illustration accompanying: Quoting Boris Cherny

Anthropic's Claude Opus 5 represents a meaningful shift in model robustness against adversarial input manipulation. According to Boris Cherny, the model demonstrates substantially improved resistance to prompt injection attacks across both internal evaluations and red-team testing, a capability gap that has plagued production LLMs since their deployment at scale. This advancement matters because prompt injection remains one of the most practical attack vectors for compromising model behavior in real-world applications, particularly in retrieval-augmented generation and multi-turn agent systems. The finding, documented in Opus 5's system card, signals that frontier labs are now prioritizing behavioral integrity alongside raw benchmark performance.

Modelwire context

Skeptical read

The claim rests almost entirely on Anthropic's own evaluations and a system card, not independent third-party audits. Prompt injection resistance is notoriously difficult to measure in a way that generalizes beyond the specific test distribution, so 'substantially improved' needs a harder number and a methodology before it carries weight.

This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It does belong to a longer-running conversation in the security research community about whether prompt injection is fundamentally an alignment problem or an input-sanitization problem, a distinction that matters because the two require different mitigations and have different failure modes in production agent pipelines.

Watch whether independent red-teamers outside Anthropic, particularly groups that have published prior prompt injection benchmarks, reproduce these resistance gains on their own test suites within the next 60 days. If the results don't replicate externally, the system card claim tells us more about eval design than about real-world robustness.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsAnthropic · Claude Opus 5 · Boris Cherny · Simon Willison

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as Quoting Boris Cherny”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Anthropic's Opus 5 shows marked resistance to prompt injection attacks · Modelwire