Inference-time rate control unlocked for autoregressive text-to-speech
Researchers have demonstrated a method to control speaking rate in autoregressive text-to-speech systems at inference time without model retraining. By identifying and manipulating a single activation axis within a decoder block, the technique enables dynamic rate adjustment while preserving speaker identity and naturalness. The approach generalizes across architectures and outperforms standard additive steering methods, particularly at extreme speeds. This work addresses a longstanding limitation in deployed TTS systems, where rate control typically requires either retraining or architectural modifications. The discovery that rate information concentrates in a low-dimensional subspace has implications for broader activation steering research and practical TTS deployment.
MentionsAutoregressive TTS · Activation steering · Decoder block analysis
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Controlling Speaking Rate in Autoregressive TTS via Activation Steering”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.