Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

Researchers benchmarked four major LLMs (GPT-4o Mini, Claude Sonnet 4, Gemini 2.5 Flash, Qwen2.5-7B) on English translation to Hausa and Fongbe, revealing stark performance gaps between typologically distinct low-resource languages and exposing reliability gaps in standard automatic metrics for African languages. The work surfaces a critical blind spot in LLM evaluation: models achieving acceptable quality on Hausa (4.0-4.5/5) fail dramatically on Fongbe (1.0-2.2/5), while BLEU and chrF++ scores often misalign with native-speaker judgment. This challenges the assumption that metric-driven model selection generalizes across linguistic families and underscores why African language coverage remains a frontier problem for production AI systems.
Modelwire context
ExplainerThe more pointed finding here is not that Fongbe scores are low, but that BLEU and chrF++ scores actively mislead practitioners by failing to correlate with native-speaker judgment, meaning teams could ship a broken Fongbe translation system while their dashboards show acceptable numbers.
This connects directly to two threads in recent coverage. The ASR corpus work on Fongbe and Hausa (published the same day) addresses the upstream data scarcity problem, but this paper shows that even when you get models to produce output, you may lack the tools to know whether that output is any good. Separately, BabelJudge, covered a day later, documents the same class of failure from a different angle: LLM-as-a-judge systems collapse in non-English contexts, which compounds the metric unreliability problem this paper identifies. Together, the three papers sketch a compounding gap: low-resource languages lack training data, lack reliable automatic metrics, and lack judges capable of evaluating them.
Watch whether any of the four benchmarked models releases updated multilingual evaluation cards that include Fongbe specifically. If none do within six months, that is a reasonable signal that low-resource African language coverage remains a compliance checkbox rather than a development priority.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT-4o Mini · Claude Sonnet 4 · Gemini 2.5 Flash · Qwen2.5-7B · COMET · BERTScore
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.