Chinese neologisms expose LLM reasoning gaps across 18 models
Researchers have exposed a significant blind spot in how leading language models handle Chinese linguistic innovation. CNeo-Bench, a curated dataset of 4,759 neologisms spanning phonetic substitution and character-based wordplay, reveals that 18 tested models struggle fundamentally with both recognizing and manipulating these expressions, with most scoring below 40% on definition tasks. The work surfaces a deeper architectural vulnerability: models can sometimes describe novel terms without grasping the generative mechanisms that produce them. This gap matters because it signals that current LLMs lack robust handling of language evolution in non-English contexts, a capability increasingly important as multilingual deployment scales.
Modelwire context
ExplainerThe paper's core insight isn't just that models fail on neologisms, but that they can sometimes generate plausible definitions without understanding the compositional rules that produce them. This decoupling between surface-level description and generative mechanism is the diagnostic signal.
This work sits alongside the August batch of LLM diagnostic papers, particularly the study on epistemic gaps in long-tail knowledge. Both expose how models systematically miss entire categories of linguistic or factual phenomena rather than failing randomly. Where that earlier work showed models omit minority viewpoints, CNeo-Bench reveals models miss the productive logic of language evolution in non-English contexts. The confidence divergence paper from the same period adds another layer: models might express high confidence on neologism definitions while internally lacking the compositional understanding to handle novel formations. Together, these papers suggest current architectures have structural blindness to dynamic, generative phenomena rather than simple knowledge gaps.
If the same 18 models show improvement on a held-out neologism test set created after training cutoff, that would confirm the gap is remediable through better training data rather than architectural. If performance remains flat, it signals the issue runs deeper than coverage and requires rethinking how models encode morphological and semantic composition in non-Latin scripts.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsCNeo-Bench · Chinese language models · LLMs
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.