
NVIDIA Audex unifies audio and text without sacrificing language ability
NVIDIA's Nemotron Labs has released Audex, a 30B-parameter multimodal LLM that unifies audio and text processing without sacrificing language performance. The model treats audio and text tokens uniformly within a single Transformer decoder, projecting audio into the text embedding space for seamless cross-modal generation. Built on a strong MoE foundation and trained on 157.4B audio tokens plus 320.5B text tokens, Audex demonstrates that audio capabilities can be added to text-dominant models through careful architecture and dataset curation. This approach matters because it sidesteps the typical capability tradeoff in multimodal scaling, potentially influencing how labs design future unified models.62

























