Speech Recognition Is Not a Solved Problem — Pavan Muddireddy
Original source
Guest
Pavankumar Reddy Muddireddy is an AI research scientist at Mistral AI.
Summary
Pavankumar Reddy Muddireddy walks through Mistral’s audio stack, starting with Voxtral Chat as an audio-input/text-output model for transcription, speaker segmentation, summarization, and question answering over long recordings. He explains the architecture as a 3B Ministral text trunk fed by an audio encoder that emits continuous embeddings directly into the decoder, avoiding Whisper-style cross-attention. The discussion then shifts to real-time ASR, where latency is explicitly conditioned as a target delay, and to diarization, hallucinations, and other failure modes that become more severe in streaming and noisy multi-speaker settings. On the generation side, Mistral’s TTS moved away from heavy discrete-codebook autoregression toward continuous latents and flow matching, with FSQ used for acoustic quantization. The episode ends on product strategy: cascades remain attractive because they are observable, adaptable, and controllable, and voice is best viewed as an augmentation layer beside a screen rather than a full replacement.