Modelwire
Subscribe

Speech translation models lose meaning by stripping disfluencies

Speech translation systems have systematized away a real problem: disfluencies like false starts and self-repairs carry semantic content that models trained on cleaned text simply drop. Researchers introduced Uh-Mazing, a multilingual benchmark pairing human translations with disfluency annotations across eight target languages, revealing that false starts and repairs account for most translation quality loss, while models tend to omit rather than mistranslate these phenomena. Crucially, inference-time decoding fixes can recover this signal without retraining, suggesting a practical path for production systems to preserve speaker intent currently lost in the pipeline.

Modelwire context

Explainer

The practical insight here is that disfluencies aren't noise to filter out but signal to preserve. Most systems treat them as artifacts of spoken language, but this work shows they carry speaker intent that text-trained models never learned to recover in the first place.

This connects directly to the multilingual retrieval and low-resource language challenges we've covered recently. The Uh-Mazing benchmark addresses a similar problem to the disentangled contrastive learning work from August 3rd: current systems degrade on non-English inputs, but here the degradation stems from a different source (disfluency handling rather than linguistic feature entanglement). Both papers identify a systematic blind spot in how we build multilingual systems. The inference-time decoding fix also echoes the practical deployment angle from the PredAct-Bench work, which stressed that production systems face real-world friction that lab evaluations miss. Here, the friction is linguistic rather than tool-related, but the pattern holds: systems trained on clean data fail when deployed against messy speech.

If Switchboard-based speech translation benchmarks start incorporating the Uh-Mazing annotations within the next six months, that signals the community is adopting this as a standard evaluation criterion. If they don't, the benchmark remains a research artifact. Also monitor whether major speech LLM providers (like those building on top of Whisper) publish disfluency-aware translation results by Q1 2027, which would indicate the inference-time fix is moving into production.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsSpeechLLMs · Switchboard · Uh-Mazing

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as The Role of Disfluencies in Speech Translation”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Speech translation models lose meaning by stripping disfluencies · Modelwire