Modelwire
Subscribe

LLMs show systematic bias when text and numbers conflict

Researchers have mapped how instruction-tuned LLMs resolve conflicts between textual and numerical evidence, revealing systematic rather than random arbitration patterns. Using a synthetic benchmark that isolates ground-truth alignment to either modality, the work independently varies source reliability, recency, and provenance to expose model decision-making. This addresses a critical gap in deployment reliability: as LLMs integrate tool outputs and structured data alongside language, understanding their evidence hierarchy becomes essential for high-stakes applications like finance and healthcare where conflicting signals are common.

Modelwire context

Explainer

The paper doesn't just show that LLMs pick text or numbers inconsistently; it reveals the decision rules they actually follow (source reliability, recency, provenance). That specificity is what makes this actionable rather than merely descriptive.

This connects directly to the MemTrapBench work from the same week, which identified how memory-augmented systems can corrupt reasoning even when inputs are accurate. Here, the problem shifts from memory fidelity to evidence hierarchy: even when both text and numbers are present and correct, the model's choice of which to trust becomes the failure point. Both papers expose gaps between input quality and downstream reliability. The autonomous driving orchestration story also resonates: just as that work treats LLMs as reasoning consultants within a larger system, this evidence arbitration research suggests LLMs need external governance over which modality to prioritize in high-stakes domains rather than autonomous decision-making.

If the researchers test their findings on real financial or medical datasets (not just synthetic benchmarks) within the next six months and show the arbitration patterns hold, that's when practitioners should start building explicit evidence-weighting layers into production systems. If the patterns collapse on real data, the work remains a curiosity about synthetic task design rather than a deployment guide.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLarge language models · Instruction-tuned models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as When Text and Numbers Disagree: Evidence Arbitration in Large Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

LLMs show systematic bias when text and numbers conflict · Modelwire