Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

Self-evaluating AI agents suffer from systematic preference collapse when using language models as judges, but multimodal systems amplify this failure dramatically. Researchers found that GPT-4o evaluating DeepSeek-chat across text and vision tasks caused a single strategy to dominate 48% of selections, triple the text-only baseline. More critically, they identified cross-modal contagion: biases learned in one modality poison strategy selection in another. This reveals a fundamental fragility in self-improving agent loops that scales with model complexity, forcing practitioners to rethink feedback architectures for multimodal systems.
Modelwire context
ExplainerThe more alarming finding isn't the collapse rate itself but the directionality: biases don't stay contained within the modality where they form. A flawed preference pattern learned while judging image tasks can corrupt text-task selection, meaning you can't isolate and fix modalities independently.
This connects directly to the misinformation propagation paper covered the same day ('Misinformation Propagation in Benign Multi-Agent Systems'), which found that false premises introduced by a single agent persist through group reasoning rather than getting corrected by debate. Both papers are describing the same underlying dynamic at different levels: feedback loops in multi-agent and self-evaluating systems amplify errors rather than damping them. The collective skill-building work in 'OpenClaw-Skill' is also relevant here, because any framework that iteratively constructs agent capabilities from self-assessed outputs inherits exactly this fragility. If the evaluator is collapsing preferences, the skill tree being built on top of those evaluations is compromised from the root.
Watch whether teams building production multimodal agent pipelines begin publishing ablations that separate evaluator modality from task modality. If that experimental design becomes standard in benchmark reporting within the next two conference cycles, this paper will have shifted how the field operationalizes evaluation hygiene.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT-4o · DeepSeek-chat · Evaluator Preference Collapse · cross-modal contagion
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.