SILICA benchmark exposes LLM agents as poor human behavior proxies
Researchers have released SILICA, an open-source benchmark that directly tests whether LLM agents replicate human behavior in social experiments or simply memorize training data. The tool runs twelve open-weight models through five environments with published human baselines, plus rule-preserving perturbations and payoff variants designed to break memorized patterns. Early results show agent-human alignment only at initialization, suggesting that apparent social dynamics may reflect dataset reproduction rather than genuine interaction. This work addresses a critical validity gap in the growing use of agent populations as computational testbeds for social science.68















