No Priors

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

Original source
Artwork for Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

Summary

In this episode, Stefano Ermon traces his path from early generative-model research at Stanford to founding Inception to commercialize diffusion-based language models. He argues that the next battleground in AI is inference, not training: autoregressive models are sequential and memory-bound at serving time, while diffusion can generate many tokens in parallel and better utilize GPUs. Ermon says Inception’s Mercury models are already in production and compare favorably with speed-optimized frontier models, but required the company to build its own serving stack because existing tooling like vLLM and SGLang does not support diffusion LLMs. He sees voice agents and other latency-sensitive applications as the first major wedge, while longer-term advantages may also include better controllability and data efficiency. The broader thesis is that efficiency, not just raw capability, will define competitive advantage in AI.

Notes