Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

Researchers audited how three leading LLMs (Claude Opus 4, GPT-4.1, Gemini 2.5 Pro) localize math word problems across seven languages spanning high-resource and under-resourced communities. The study reveals whether models preserve cultural specificity or collapse diverse contexts into generic translations, exposing systematic biases in how AI systems handle cultural entities at scale. This matters for educational deployment: if models strip local context or over-generalize, personalized learning tools risk erasing cultural representation while appearing neutral. The findings signal a gap between multilingual capability claims and actual cultural fidelity in production systems.
Modelwire context
ExplainerThe study's framing around math word problems is deliberate and pointed: arithmetic contexts are often assumed to be culturally neutral, which makes them an ideal stress test for exposing invisible normalization. The real finding isn't that models fail at translation, it's that they can pass surface-level fluency checks while still erasing the cultural specificity that makes localized content meaningful.
This connects directly to the June 9th paper on 'Measuring Human Value Expression in Social Media Texts,' which tackled a parallel problem: LLMs annotating subjective cultural constructs at scale without reliable grounding in the communities those constructs belong to. Both papers are probing the same structural gap, which is the distance between what a model can process linguistically and what it actually represents culturally. That earlier work focused on annotation fidelity for values; this one focuses on generative fidelity for context. Together they suggest a pattern worth tracking: multilingual and culturally-aware are being conflated in capability claims, and researchers are now systematically building the audit infrastructure to separate them.
Watch whether any of the three evaluated model providers (Anthropic, OpenAI, Google) cite this audit in future multilingual benchmark disclosures or update their educational product documentation to distinguish linguistic coverage from cultural representation. Silence on that front would itself be informative.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsClaude Opus 4 · GPT-4.1 · Gemini 2.5 Pro
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.