Code-switched speech exposes hidden failures in top ASR models
Modern speech recognition systems mask critical failures on code-switched speech, a finding that exposes a blind spot in how the field measures progress. Researchers evaluated eleven state-of-the-art ASR and audio language models on English-Yoruba mixed-language utterances, revealing that aggregate word error rate obscures systematic breakdowns at language boundaries and with diacritics. The best performer by conventional metrics actually struggled most with switch points, suggesting that benchmark dominance on monolingual data provides false confidence for real-world deployment in multilingual contexts. This work signals that evaluation methodology itself needs restructuring to catch failure modes invisible to standard metrics.
Modelwire context
ExplainerThe critical finding isn't just that models fail on code-switched speech, but that the failure is invisible to standard metrics. A model ranked first by word error rate can rank last when you measure performance specifically at language switch points, suggesting the field has been optimizing for the wrong target.
This work sits directly alongside IndicTriMix (token-level language identification in code-mixed text) and Nuha-Speech (Arabic speech-LLM infrastructure), which both treat multilingual handling as a distinct problem requiring dedicated benchmarking rather than hoping monolingual scaling solves it. Where those papers build infrastructure for underrepresented languages, this one exposes why generic evaluation frameworks mask failures in those exact contexts. The pattern across all three is the same: real-world linguistic diversity demands measurement and datasets tailored to that diversity, not borrowed from monolingual baselines.
If major ASR vendors (Whisper, Conformer-based systems) publish switch-aware error breakdowns on their own benchmarks within six months, that signals the field is adopting the metric. If they don't, and code-switched speech remains unmeasured in standard leaderboards through 2027, this paper will have identified a problem the industry isn't ready to fix.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsEnglish-Yoruba code-switched speech · ASR models · audio language models · word error rate · switch entry token error rate
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.