
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
Researchers have identified a critical vulnerability in deployed language models: stealth biases that favor specific entities or viewpoints while remaining invisible to standard audits. The threat emerges when bad actors embed preferential signals into soft logit distributions during model distillation, making detection nearly impossible without prior knowledge of the target bias. This work exposes a fundamental asymmetry in AI safety: defenders cannot reliably catch hidden steering attacks without knowing what to look for, raising urgent questions about supply chain integrity and the adequacy of current model inspection techniques for high-stakes deployments.68




























