← Back to blog

LLMs Now Generate Tokens in Parallel Without Losing Accuracy

Based on research by Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi

Large Language Models are powerful but painfully slow, forced to generate text one token at a time like a reader tracing each word with a finger. This sequential bottleneck limits their speed, but researchers have now introduced a method that breaks this chain without sacrificing accuracy. By combining the reliability of autoregressive models with the parallel power of diffusion, they have created a new class of models that generate multiple tokens simultaneously.

The core innovation lies in decoupling the model into two parts. The main autoregressive weights handle the standard next-token prediction, while lightweight diffusion weights are trained to generate several tokens in parallel. This diffusion training is added via a simple distillation process that adds negligible overhead to existing pipelines. The result is a system that draws from the same distribution as the original model but does so in a burst, effectively parallelizing the generation process.

This approach solves a major conflict in current acceleration techniques. Unlike speculative decoding, it requires no separate draft model to guess the next steps. Unlike other diffusion-based LLMs, it does not degrade the quality of the underlying autoregressive model. The researchers call these new models Uno and introduce a sampler family that enables lossless acceleration. This means you get faster speeds without the typical trade-off of reduced coherence or accuracy.

The performance gains are substantial. Uno models achieve higher throughput than leading speculative-decoding methods across all batch sizes and deliver up to three times the speed of the base autoregressive model. Notably, their eight-billion-parameter Uno model outperforms larger twenty-six-billion-parameter diffusion models and proprietary systems in coding, tool use, and reasoning benchmarks. This breakthrough allows developers to build faster, more efficient AI systems by simply augmenting existing open-weight models with these new diffusion capabilities.

Source: arXiv:2609.04010

This post was generated by staik AI based on the academic publication above.