What happened after 2,000 people tried to hack my AI assistant
Source published ·Modelwire updated
Original coverage: Simon Willison ↗·How Modelwire adds context

The development
Fernando Irarrázaval's public red-teaming experiment exposed a critical gap between prompt-injection resilience claims and real-world robustness. Over 6,000 adversarial attempts against an Opus 4.6 instance with explicit anti-injection rules failed to extract secrets, suggesting either that modern LLM safeguards are holding under sustained attack or that the test's constraints were too narrow to surface vulnerabilities. The finding matters because it challenges both the doomsday narrative around prompt injection and the assumption that simple rule-based defenses suffice, forcing the field to recalibrate expectations around LLM security posture at scale.
Modelwire’s AI-generated summary of coverage from Simon Willison.
Modelwire analysis
Skeptical readOur AI-generated reading of the wider context and the next developments to watch.
The experiment's headline number, 6,000 adversarial attempts with zero secret extraction, obscures a more important question: what attack categories were actually attempted, and were the most sophisticated multi-turn, context-manipulation techniques represented in that pool of 2,000 participants, most of whom were likely casual rather than expert adversaries.
This is largely disconnected from recent activity in our archive, as we have no prior coverage of prompt injection research, red-teaming methodology, or Anthropic's Opus model line to anchor it against. It belongs to a broader conversation in the security research community about whether public red-teaming exercises produce generalizable findings or simply measure the resilience of a specific configuration against a self-selected, non-expert crowd. That distinction matters enormously when interpreting the result.
Watch whether Irarrázaval or an independent party publishes a breakdown of attack taxonomy from the attempt logs. If the dataset shows fewer than 5% of attempts used multi-turn or indirect injection strategies, the robustness claim weakens considerably and the experiment tells us more about attacker skill distribution than about model defenses.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsFernando Irarrázaval · OpenClaw · Anthropic Claude Opus 4.6 · hackmyclaw.com
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.