
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
Researchers have formalized a method to reverse-engineer the pretraining data composition of LLMs by analyzing only their generated outputs. LLMSurgeon treats this as an inverse problem, using calibrated confusion matrices to estimate domain-level distributions across a predefined taxonomy without access to training corpora. This addresses a critical transparency gap: most frontier labs keep data mixtures proprietary, blocking external audits of model provenance and potential contamination. For practitioners and safety researchers, the ability to forensically decompose a model's training diet from behavior alone reshapes accountability and competitive benchmarking, especially as data provenance becomes a regulatory and reputational concern.62




























