Moshi-based model adds instruction-driven control to full-duplex speech
Full-duplex speech models enable natural back-and-forth conversation but have lacked fine-grained control over conversational attributes. SteerDuplex addresses this gap by extending Moshi with instruction-following capabilities across tone, persona, speaking rate, and voice style. The work combines synthetic dialogue training with two-stage reinforcement learning using hybrid rewards, establishing a taxonomy of steerability dimensions that were previously unmeasured in production systems. This matters because controllable speech interaction is foundational for enterprise dialogue systems, accessibility tools, and personalized voice assistants where users need reliable behavioral adaptation without retraining.
Modelwire context
ExplainerSteerDuplex's actual contribution is narrower than it sounds: it's not a new model, but a training methodology that adds instruction-following to Moshi. The key technical move is using hybrid rewards (combining multiple objectives) during RL to make tone, persona, and voice style controllable without degrading base conversation quality.
This lands in the middle of a coordinated push on full-duplex dialogue naturalness. DuplexDrama (released the same day) provided the messy, overlapping training data that systems like Moshi need. SteerDuplex now addresses the inverse problem: once you have natural dialogue, how do you steer it reliably? The 'Continue, Adapt, or Yield' framework from the same date identified that agents struggle with mid-turn adaptation; SteerDuplex's persona and tone controls are preconditions for that kind of nuanced behavioral adjustment. Together, these papers suggest the field is moving from 'can we make full-duplex work at all' to 'can we make it controllable and contextually appropriate.'
If SteerDuplex's steerability taxonomy (the dimensions it measures) gets adopted by other full-duplex models or benchmarks within the next six months, that signals real standardization. If instead each new model invents its own control dimensions, the work remains a one-off engineering contribution rather than a shared framework.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSteerDuplex · Moshi · arXiv
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “SteerDuplex: Steerable Duplex Speech Dialogue Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.