Modelwire
Subscribe

SILICA benchmark exposes LLM agents as poor human behavior proxies

Researchers have released SILICA, an open-source benchmark that directly tests whether LLM agents replicate human behavior in social experiments or simply memorize training data. The tool runs twelve open-weight models through five environments with published human baselines, plus rule-preserving perturbations and payoff variants designed to break memorized patterns. Early results show agent-human alignment only at initialization, suggesting that apparent social dynamics may reflect dataset reproduction rather than genuine interaction. This work addresses a critical validity gap in the growing use of agent populations as computational testbeds for social science.

Modelwire context

Skeptical read

SILICA's real contribution isn't proving agents memorize social data (that's expected). It's the perturbation strategy: by tweaking payoffs and rule sets, the benchmark attempts to distinguish genuine behavioral learning from dataset reproduction. But the paper doesn't clarify whether agents trained on diverse social datasets would show better generalization, or whether the problem is fundamental to how LLMs encode social reasoning.

This connects directly to the confidence-divergence work from August and the persona-dialogue study. Both exposed gaps between what models claim to do and what they actually execute. SILICA extends that pattern: agents appear to engage in social dynamics, but the internal mechanism (memorization vs. reasoning) remains opaque. The PersonaForge and CultureConverse papers from the same week also highlight how multi-turn interaction exposes failures that single-shot evaluation misses, reinforcing SILICA's core finding that static benchmarks underestimate agent brittleness.

If the same twelve models show improved agent-human alignment when fine-tuned on SILICA's perturbation variants (rather than original environments), that suggests the gap is remediable and points toward social reasoning as learnable. If alignment stays flat across retraining, the memorization hypothesis holds and agent-based social simulation becomes fundamentally limited for social science.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsSILICA · arXiv

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Benchmarking large language model agent societies against human behavioural distributions”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

SILICA benchmark exposes LLM agents as poor human behavior proxies · Modelwire