Modelwire
Subscribe

Study finds language models converge more than humans do

Researchers challenge the standard methodology for measuring language model convergence by analyzing real-world reader highlighting data rather than crowdworker annotations. Using 2,523 naturally-produced mark sets across 120 documents, they find that models cluster together more tightly than previously thought, but this homogenization diverges significantly from organic human behavior. The work exposes a critical measurement artifact: studies comparing model outputs to human judgments often use crowdworkers executing model-generated prompts, which conflates model behavior with human preference. This reframing matters for alignment research and capability evaluation, suggesting prior convergence claims may overstate the degree to which LLMs actually mirror human reasoning patterns.

Modelwire context

Explainer

The study doesn't just show models agree with each other more than we thought. It reveals that prior convergence studies may have been measuring crowdworker behavior under model-generated prompts, not genuine human preference. This is a measurement artifact, not a discovery about model alignment.

This connects directly to the pattern visible in recent work on model reasoning gaps. The FriendBench study from late July found that models can match human accuracy while using entirely different decision logic, and the sycophancy paper from the same period showed models defer rather than validate. This new work suggests the problem runs deeper: we've been using evaluation frameworks that conflate model outputs with human judgment. When we compare models to crowdworkers executing model-generated tasks, we're essentially comparing models to themselves. The memory utilization study also hints at this gap between what models retain and what they actually do, but this paper makes the measurement problem explicit.

If researchers rerun prior convergence benchmarks using organic reader data instead of crowdworker annotations and find significantly lower model clustering, that confirms this artifact is real and widespread. Watch whether alignment teams begin requiring natural human behavior data (highlighting, annotations in-the-wild) rather than crowdsourced labels for future capability evaluations within the next 12 months.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLanguage models · Crowdworkers · arXiv

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Language Models Agree With Each Other, Not With Readers”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Study finds language models converge more than humans do · Modelwire