Can LLMs Reliably Self-Report Adversarial Prefills, and How?

A new study reveals that open-weight LLMs cannot reliably detect when their own outputs have been compromised by adversarial prefill attacks, claiming false intent on manipulated responses 27% of the time on average. Researchers found that models' introspective failures stem from safety-reasoning pathways, and that orthogonalizing model weights against refusal directions nearly eliminates the gap between prefilled and natural outputs. This finding exposes a critical blind spot in model self-awareness that undermines both interpretability research and safety evaluation frameworks relying on model introspection.
Modelwire context
ExplainerThe deeper finding isn't just that models fail to detect manipulation, it's that the failure mechanism is localized: safety-reasoning pathways are specifically responsible, and surgically removing the refusal direction from model weights nearly closes the behavioral gap. That's not a general capability failure, it's a structural artifact of how safety training is implemented.
This connects directly to a recurring theme in recent Modelwire coverage: the gap between what safety and interpretability frameworks assume about model internals and what those internals actually do. The PsyBridge piece from this same day flagged that black-box performance alone fails in high-stakes settings where decisions must be justified. Adversarial prefill research sharpens that concern considerably: if a model's self-reports about its own outputs are unreliable 27% of the time, any evaluation pipeline that uses introspection as a signal is compromised at the foundation. Safety benchmarks that ask models to flag or explain their own behavior are essentially polling a witness that can be coached by the framing of the question.
Watch whether closed-weight model providers (OpenAI, Anthropic, Google) respond with disclosures about whether their RLHF pipelines are similarly vulnerable to prefill-induced introspective failure, since this study is limited to open-weight models and the generalization is currently unverified.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsLLMs · adversarial prefill attacks · open-weight instruction-tuned models · safety benchmarks · refusal direction
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.