LLM value measurement methods show inconsistent preference signals across tasks
STONIC exposes a fundamental crack in how researchers measure LLM values and preferences. By testing 35 model configurations across 5,144 scenarios, the study reveals that questionnaires, pairwise choices, and spontaneous text generation do not measure the same underlying preference, contradicting a core assumption in value alignment work. Most critically, models consistently favor their own prior outputs over alternatives, and choice patterns shift based on option ordering, suggesting that current value profiling methods may be capturing behavioral artifacts rather than stable preferences. This finding directly challenges the validity of value datasets used to train and evaluate alignment in production systems.
Modelwire context
Skeptical readThe paper's real contribution is narrower than the summary suggests: it documents that measurement method matters, which is unsurprising in social science. What's missing is whether STONIC actually proves preferences are unstable or merely shows that questionnaires, rankings, and generation tap different behavioral signals (a known problem in preference elicitation). The 'behavioral artifacts' claim needs scrutiny.
This connects directly to the weird generalization work from earlier this month, which found that measurement sensitivity to question selection significantly impacts evaluation reliability. Both papers expose how fragile our confidence in LLM preference measurement actually is. However, STONIC goes further by questioning whether stable preferences exist at all, whereas the WG research treated measurement variance as a calibration problem. The gap between these two framings matters: one says 'we're measuring wrong', the other says 'there may be nothing stable to measure'.
If STONIC's authors release ablations showing that option ordering effects persist when controlling for presentation bias (e.g., randomizing position across multiple trials per model), that strengthens the instability claim. If the ordering effects disappear under controlled presentation, the finding collapses to a measurement artifact rather than evidence of preference incoherence. Watch for follow-up work testing whether the same models show consistent preference rankings when measurement method is held constant.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSTONIC
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “STONIC: A Layered Measurement Contract for LLM Value Profiling”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.