LLMs show under 6% agreement with expert physics literature searches
A controlled study comparing mid-2025 LLMs against expert physicists on literature review tasks reveals a critical gap in AI research assistance. ChatGPT-4o, ChatGPT Deep Research, and Gemini showed less than 6% overlap with human-selected references, signaling that current models cannot yet independently replicate expert-level search strategies. The finding reframes LLM utility in scientific workflows from autonomous replacement to complementary tool, with implications for how researchers should architect AI-assisted discovery pipelines and where model training must improve to bridge domain-specific retrieval gaps.
Modelwire context
Skeptical readThe paper doesn't clarify whether the 6% overlap reflects genuine retrieval gaps or misalignment between how the study prompted models versus how domain experts naturally search. If researchers typically use LLMs to expand or cross-check curated lists rather than bootstrap searches from scratch, this benchmark may measure the wrong task.
This connects directly to the WorkSurface-Bench finding from late July, which identified that enterprise agents struggle with knowledge source routing before retrieval even begins. Both studies reveal that current models lack the meta-cognitive layer to identify what kind of search strategy a task requires. The physics literature review problem is a domain-specific instance of the same routing failure: models don't know whether to prioritize citation networks, keyword matching, or temporal relevance the way expert physicists do implicitly.
If the researchers release ablation data showing performance when given explicit search strategy hints (e.g., 'prioritize recent papers citing X'), overlap should jump significantly. If it doesn't, that confirms the gap is semantic understanding, not just retrieval. Watch whether follow-up work tests whether fine-tuning on physics-specific search logs improves overlap above 15 percent within six months.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsChatGPT-4o · ChatGPT Deep Research · Gemini · OpenAI · Google
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.