Self-harm representations crystallize in final model layers across architectures
Researchers mapped where language models encode self-harm content across their internal layers, finding that harmful semantic information concentrates in the final 7% of network depth across multiple architectures. This mechanistic insight has direct implications for safety interventions: by identifying the specific layers where harmful representations crystallize, practitioners can design more targeted detection systems and potentially intervene earlier in the model's computation pipeline. The cross-architecture validation suggests these findings generalize beyond individual model families, making this a foundational step toward interpretable content moderation at scale.62















