
Survey maps efficiency bottlenecks in parallel diffusion language models
Diffusion-based language models promise parallel generation advantages over sequential autoregressive architectures, but converting theoretical speedups into real-world deployment requires careful system design. This survey maps the emerging landscape of acceleration techniques spanning algorithm optimization, hardware architecture, and inference infrastructure, while establishing a latency decomposition framework to isolate the true sources of efficiency gains. The work addresses a critical gap in benchmarking rigor, where end-to-end performance metrics often obscure which optimizations actually matter in production. For practitioners evaluating dLLM viability, this structured analysis clarifies where engineering effort yields returns versus where trade-offs remain unresolved.58






















