Full-duplex speech agents vulnerable to imperceptible audio attacks, defense proposed
Real-time speech dialogue systems now face a critical vulnerability: imperceptible audio attacks that exploit the acoustic masking properties of human hearing. Researchers demonstrated that full-duplex agents like Moshi can be hijacked, silenced, or jailbroken through adversarial perturbations hidden beneath the psychoacoustic threshold, with success rates exceeding 90%. The proposed defense, psychoacoustically aligned latent smoothing, injects structured noise at the quantized latent layer to disrupt attacks while preserving speech quality. This work signals that multimodal conversational AI requires fundamentally different robustness assumptions than text systems, forcing practitioners to integrate psychoacoustics into threat modeling for production dialogue agents.
Modelwire context
ExplainerThe critical insight here isn't just that full-duplex speech models are vulnerable to adversarial audio, but that the attack surface differs fundamentally from text systems because it exploits human auditory perception itself. Defenses must account for what humans can't hear, not just what models can't parse.
This work sits at the intersection of two recent findings in our coverage. Like the demographic bias audit from late September, it exposes a modality-specific failure mode in speech systems that text-only robustness assumptions miss entirely. But it also connects to the multimodal optimization challenge covered in VCMM, where different modalities require specialized handling during training. Here, the defense operates at the latent layer rather than raw audio, suggesting that robustness, like training efficiency, demands modality-aware architecture choices rather than generic fixes.
If Moshi or other commercial full-duplex agents ship with psychoacoustically aligned smoothing in production within six months, that signals the threat is real enough to warrant deployment overhead. If no major speech dialogue system adopts this by mid-2027, it suggests either the 90% attack success rate doesn't replicate on real-world deployments or the latency cost is prohibitive.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsMoshi · psychoacoustically aligned latent smoothing · full-duplex speech-to-speech dialogue models
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Psychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.