Spiking the training data to correct for test set contamination
Researchers propose a novel approach to correcting inflated test scores caused by data leakage, a persistent problem in model evaluation. Rather than only detecting contamination, the method intentionally spikes training data with known test examples to calibrate memorization predictors, enabling statistical adjustment of benchmark results. The work introduces Hubble models as a simulation framework with paired contaminated and clean variants to validate correction estimators. This addresses a critical gap in ML rigor: while test set contamination is widely acknowledged, principled correction methods remain rare. The technique could reshape how labs validate model performance and report benchmark claims, particularly as model scale makes accidental data leakage increasingly likely.62




























