
Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?
Vision-language models exhibit a critical failure mode: they generate high-confidence predictions on visually ambiguous inputs while standard uncertainty quantification methods fail to detect it. Researchers show that entropy-based approaches like Semantic Entropy underestimate uncertainty because overconfident visual embeddings suppress output diversity during decoding. Perturbation-based alternatives designed to probe robustness instead conflate textual sensitivity with visual understanding, masking the core problem. This work exposes a fundamental gap in how we measure VLM reliability, with direct implications for deployment in safety-critical domains where false confidence on ambiguous visual inputs poses real risk.62




























