ConlangBench reveals LLM reliance on lexical shortcuts over compositional learning
Researchers have built ConlangBench, a 21-million-sentence parallel corpus spanning 21 constructed languages, to probe how LLMs acquire and generalize linguistic structure beyond natural language. The benchmark reveals that models excel on a posteriori conlangs, whose vocabularies anchor to existing languages, suggesting LLMs exploit surface-level lexical patterns rather than learning deep compositional rules. This work matters because it exposes a fundamental gap in how current models handle systematic language design, with implications for transfer learning, few-shot adaptation, and whether scaling alone can produce genuine linguistic reasoning.
Modelwire context
ExplainerThe key insight isn't just that models fail on constructed languages, but that they fail *differently* on a posteriori conlangs (those anchored to natural language vocabularies) versus a priori ones (fully synthetic). This asymmetry suggests models are pattern-matching on surface lexical overlap rather than learning the underlying compositional structure the benchmark was designed to test.
This joins a cluster of recent work using specialized benchmarks to expose gaps between apparent capability and actual reasoning. Like TreeProbe's measurement of cultural bias through native epistemic structures and ChronoLens's use of frozen models as linguistic instruments, ConlangBench treats the benchmark itself as a diagnostic tool rather than just a leaderboard. The common thread: existing benchmarks (whether multilingual corpora or general knowledge tests) may conflate surface performance with genuine understanding. ConlangBench isolates compositional reasoning in a way natural language benchmarks cannot, similar to how FinHardBench isolated latency constraints that code generation benchmarks miss.
If ConlangBench performance correlates with downstream transfer learning success on low-resource natural languages (languages with <1M training tokens), that validates the claim that a posteriori/a priori distinction predicts real generalization. If performance remains decoupled from transfer metrics, the benchmark measures an artifact of the conlang design rather than a fundamental model limitation.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsConlangBench · Esperanto
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.