KIGALI — Inception Labs, a generative AI startup founded by Stanford, UCLA, and Cornell faculty, has launched Mercury 2.5, the first commercially available diffusion large language model to reach production-scale performance.
The model, released on September 8, 2026, generates 1,107 tokens per second on NVIDIA GPUs—more than 3x faster than traditional auto-regressive LLMs—while consuming less than half the compute cost per token. This performance gain comes without sacrificing quality: Mercury 2.5's intelligence score, as measured by independent benchmarks, jumped 40% over Mercury 2.
Why diffusion matters
Traditional LLMs generate text one token at a time, like a person typing a letter. Diffusion models, by contrast, predict thousands of tokens in parallel through an iterative refinement process similar to how image diffusion works. This architectural shift enables parallel output, which is the key to both speed and cost efficiency.
"Diffusion is not just about images," said Volodymyr Kuleshov, co-founder and CTO of Inception Labs. "When you apply the same principles to language, you get a model that thinks in parallel. That's how you get to 1,100 tokens per second while cutting costs by more than 50%."
Technical specifications
Mercury 2.5 is built on Inception's proprietary diffusion transformer architecture. Key capabilities include:
-
260K context window: The model can process documents up to 260,000 tokens long, enabling analysis of entire codebases, research papers, or legal contracts in a single pass.
-
Fine-grained control: Users can constrain outputs to specific schemas or semantic requirements, making the model suitable for structured data extraction and API generation.
-
Multimodal readiness: While the current release is text-only, the diffusion paradigm unifies language generation with audio, image, and video capabilities in a single model family.
The model is now available through Microsoft Azure AI, making it accessible to enterprise customers without requiring specialized infrastructure.
Performance benchmarks
Inception Labs reported testing results on NVIDIA H100 GPUs, showing inference speeds of 1,107 tokens per second. Independent analysts have confirmed the speed advantage, noting that auto-regressive models typically cap out at 300-400 tokens per second on comparable hardware.
"We saw a 2.8x speedup in real-world API tests when comparing Mercury 2.5 to the best auto-regressive models we evaluated," said one developer at a Fortune 500 company deploying the model for document summarization.
The cost advantage stems from fewer forward passes. Auto-regressive models must run one forward pass per token, while diffusion models can generate multiple tokens in parallel, reducing the total number of GPU operations needed.
Team and background
Inception Labs was founded by three academic researchers: Stefano Ermon, a Stanford professor and co-inventor of diffusion models, flash attention, and Direct Preference Optimization (DPO); Aditya Grover, a UCLA professor known for node2vec and decision transformers; and Kuleshov, formerly co-founder and CTO of Afresh Technologies.
The engineering team includes researchers from Google DeepMind, Meta AI, Microsoft AI, and OpenAI. The company is already deploying large-scale diffusion models at Fortune 500 firms.
Implications for enterprise adoption
The speed and cost improvements could accelerate AI adoption in latency-sensitive applications. For example, customer service chatbots, real-time translation services, and live document analysis tools could all benefit from the higher throughput.
Industry analysts note that diffusion LLMs may also make multimodal applications more feasible, since the same architecture can handle text, audio, and visual data. Inception Labs has hinted that audio and video models using the same diffusion paradigm will follow in coming months.
Looking ahead
Inception Labs is currently scaling its deployment infrastructure to meet enterprise demand. The company plans to release additional model variants optimized for specialized tasks including code generation, reasoning, and voice agent applications.
"Mercury 2.5 is just the beginning," Kuleshov said. "We're going to show that diffusion can do everything auto-regressive models do, but faster and cheaper. That's the goal."
As of mid-September 2026, Mercury 2.5 remains the only production-grade diffusion LLM with verified high-speed performance. The question now is whether other providers will follow suit with their own diffusion architectures, or whether Inception Labs can maintain its technical lead.
Sources
- Inception Labs, September 8, 2026 —
- BenchLM.ai, September 10, 2026 —
- Shattered.io, September 8, 2026 —
FIRAT Editorial Team
Research Contributor · Foresight Institute of Research and Translation
FIRAT Editorial Team contributes to FIRAT's mission of generating evidence-based research and translating scientific breakthroughs into sustainable African development.

