Closing CLIP's modality gap can harm zero-shot accuracy
A new analysis of CLIP's cross-modal alignment reveals a counterintuitive failure mode: shrinking the image-text representation gap does not guarantee better zero-shot classification. Researchers show that correcting modality misalignment can paradoxically degrade performance by concentrating predictions onto a narrow set of classes, a phenomenon termed prediction-level hubness. This challenges the conventional wisdom that tighter multimodal alignment uniformly improves downstream tasks, forcing practitioners to reconsider how alignment metrics map to actual decision quality in vision-language systems.62





