UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

UrduMMLU addresses a critical gap in multilingual AI evaluation by introducing the first large-scale, native-source benchmark for Urdu language understanding. With 26,431 questions spanning 26 subjects drawn from authentic educational materials rather than translations, the benchmark enables rigorous assessment of how well LLMs perform in a language spoken by 230 million people. Testing 30 models across English and Urdu prompts reveals significant performance disparities, signaling that current benchmarking practices may mask real-world capability gaps in non-English contexts. This work underscores why language-specific, culturally grounded evaluation is essential as AI deployment expands globally.
Modelwire context
ExplainerThe critical methodological choice here is sourcing questions from authentic Pakistani and Indian educational materials rather than translating existing English benchmarks, which eliminates the translation artifacts that typically inflate scores and obscure genuine comprehension failures. That design decision is what makes the performance gaps UrduMMLU surfaces credible rather than noise.
This fits directly into a pattern Modelwire has been tracking across multiple recent papers. K-BrowseComp (June 1) showed frontier models dropping 30-45 percentage points on Korean web tasks compared to English equivalents, and UrduMMLU is now producing analogous disparity findings for Urdu at a much larger subject-coverage scale. The FRANZ audit framework covered here on June 1 adds another layer: even where models score adequately on factual benchmarks, communicative framing in non-English contexts remains unaudited. Together these papers suggest the evaluation gap for non-English languages is structural, not incidental, and that no single benchmark type captures it fully.
Watch whether the teams behind leading multilingual models (Gemini-3 is already tested here) publish targeted fine-tuning or retrieval-augmented results on UrduMMLU within the next six months. Sustained low scores with no response from major labs would confirm that Urdu remains a genuine capability blind spot rather than a fixable data gap.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsUrduMMLU · Gemini-3 · MMLU
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.