Romanization outperforms native scripts in multilingual model pretraining
Researchers systematically evaluated how multilingual language models handle languages with different writing systems, comparing romanization, IPA transcription, and native orthography during pretraining. Testing across three model scales (467M to 1B parameters) and eight typologically diverse languages, romanized inputs consistently outperformed alternatives on downstream tasks spanning both seen and unseen languages. This finding challenges the assumption that native scripts are optimal for cross-lingual transfer and has immediate implications for practitioners building multilingual systems, particularly for low-resource language pairs where script mismatch currently limits knowledge sharing.
Modelwire context
ExplainerThe paper doesn't just show romanization works; it demonstrates that script-agnostic pretraining creates better cross-lingual transfer than preserving native orthography. This matters because it suggests the model learns language structure more robustly when decoupled from writing system idiosyncrasies.
This connects directly to the Baniwa ASR work from the same day, which adapted Whisper (a multilingual model trained on high-resource languages) to an endangered language with minimal data. That study showed foundation models can be repurposed for under-resourced languages, but it didn't address the script bottleneck. This romanization finding removes one barrier to that repurposing: practitioners building systems for low-resource language pairs no longer need to assume native orthography is a prerequisite for knowledge transfer. Together, these papers suggest a clearer pathway for extending multilingual infrastructure to linguistic margins.
If teams building ASR or translation systems for low-resource language pairs adopt romanized pretraining in the next 12 months and report downstream task improvements matching or exceeding native-script baselines, that confirms the finding generalizes beyond the controlled eight-language test set. If adoption stalls or shows mixed results on real-world data, the result may be an artifact of the evaluation setup.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.