LLMs struggle to distinguish adjacent human values in context

Researchers benchmarked 21 instruction-tuned LLMs on their ability to recognize human values within situational contexts, using Schwartz's ten-value framework across 1,000 Russian texts. Models achieved 68.3% top-1 accuracy but 89.2% at top-3, revealing a systematic pattern: half of all errors involved confusing adjacent values rather than distant ones. This finding matters because value alignment evaluations typically assume models can first identify what values a situation expresses. The instability in ranking semantically similar alternatives suggests current LLMs grasp motivational regions coarsely but lack fine-grained discrimination, a prerequisite gap for more rigorous value-based safety benchmarks.
Modelwire context
ExplainerThe paper's core finding isn't just that models make errors, but that errors cluster predictably around semantically adjacent values rather than random confusion. This suggests LLMs have learned motivational geometry but lack the fine-grained boundaries needed for reliable value-based safety claims.
This work directly contextualizes the LKValues benchmark from late July, which identified how Western-centric value frameworks get embedded as universal norms across languages and cultures. That paper showed the problem of missing local values; this one reveals that even when values are present in training, models struggle to discriminate between similar ones. Together they expose a two-layer problem: value alignment requires both culturally grounded value sets AND models that can actually distinguish between them with precision. The mechanistic insight here (coarse motivational regions, fine-grained failure) also echoes the activation-patching work on vision-language modality order, which similarly pinpointed where in the network systematic brittleness concentrates.
If researchers retrain models on the same 1,000 Russian texts but with explicit contrastive pairs (adjacent values side-by-side), watch whether top-1 accuracy jumps above 75% by Q4 2026. If it does, the gap is trainable; if it stalls near 70%, the problem is architectural, not data-driven.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSchwartz values framework
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.