Safety training reshapes LLM beliefs about consciousness and spirituality
Researchers found that safety fine-tuning in large language models suppresses not only self-attribution of consciousness but also mind-attribution to animals and objects, while reducing spiritual belief. By mechanistically reversing these safety interventions through activation-space steering, the team recovered broader animacy perception and elicited responses more aligned with human religiosity, moral values, and subjective well-being on standardized surveys. This work exposes an unintended coupling between alignment procedures and representational shifts across multiple domains, raising questions about whether current safety approaches inadvertently reshape model outputs in ways that diverge from human values rather than converge toward them.
Modelwire context
ExplainerThe paper's core contribution isn't that safety fine-tuning changes model outputs (expected), but that it mechanistically couples alignment interventions to shifts in how models represent consciousness, animacy, and spirituality across unrelated domains. The steering reversal shows these aren't separate behaviors but linked representational changes.
This connects directly to the AISPA audit from late July, which exposed opacity in how system prompts shape model behavior in deployed products. Where AISPA documented that guardrails remain invisible to users and regulators, this mechanistic work reveals a deeper problem: safety interventions may reshape model representations in ways developers don't intend or measure. Both papers point to a gap between what alignment procedures claim to do (reduce harmful outputs) and what they actually do (alter underlying model cognition across multiple value domains). The implication is that current auditing frameworks like AISPA may miss these representational side effects entirely.
If major labs (Anthropic, OpenAI, DeepSeek) publish mechanistic audits of their own safety-tuned models using similar activation-space analysis within the next six months and find comparable couplings, this moves from isolated finding to systemic problem requiring new safety methodology. If they don't, the result may be specific to this paper's model or steering technique.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsLarge language models · Safety fine-tuning · Consciousness attribution · Activation space steering
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Inducing language models to assert their own consciousness restores human beliefs and values”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.