Modelwire
Subscribe

LLM preference judgments fail basic consistency tests

A new study challenges a core assumption in preference-learning pipelines: that LLM-generated preference judgments are internally consistent. Researchers measured whether cardinal judgments like willingness-to-pay satisfy basic transitivity and equivalence properties, developing statistical tests to quantify inconsistency. The finding matters because many AI systems now infer utility functions from natural-language LLM outputs to guide decision-making. If those outputs violate fundamental axioms of rational preference, downstream actions based on estimated utilities may be misaligned with actual user intent, raising questions about the reliability of LLM-as-preference-oracle architectures in high-stakes applications.

Modelwire context

Explainer

The study doesn't just show LLMs make inconsistent judgments; it quantifies the violation of basic rational axioms (transitivity, equivalence) in cardinal preference outputs. This is distinct from general hallucination or factual errors: it's structural irrationality in the preference signal itself.

This finding sits alongside a cluster of August papers exposing hidden brittleness in LLM-as-oracle systems. The 'Judge, Retrieve, or Abstain' framework from the same week proposes uncertainty quantification to make LLM judges safer; this preference inconsistency paper suggests the problem runs deeper than missing confidence scores. Similarly, the latent-structures study showed that LLMs and humans solve problems via different mechanisms, implying that inferring human utility from LLM outputs may conflate surface agreement with fundamentally misaligned reasoning. And the belief-handling paper revealed that LLM outputs fluctuate based on phrasing rather than stable capability. Together, these suggest that treating LLM outputs as direct proxies for human intent (whether for judging, preferring, or reasoning) systematically underestimates model unreliability.

If preference-learning pipelines that incorporate explicit transitivity constraints (e.g., via constraint satisfaction or post-hoc reconciliation) show measurable improvement in downstream task alignment within the next six months, that confirms inconsistency is a fixable bottleneck. If they don't, it signals the problem is deeper than logical structure and points toward fundamental misalignment between LLM outputs and human utility.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLLM

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as LLM-Derived Preference Judgments Are Not Self-Consistent”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

LLM preference judgments fail basic consistency tests · Modelwire