How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen
Original source
Summary
Shawn Wen makes the case that voice is the next major interface, but also the hardest one because it adds time, interruptions, and physical-world variability to the problem of AI. He says enterprise buyers want a contradictory mix: human-like capability with robot-like controllability, which pushes PolyAI toward branded voices, citations, audit trails, and strict guardrails. The company moved away from cascaded ASR/TTS pipelines toward an audio-native architecture that predicts turn-taking, then emits text/tool calls, then citations and transcripts for auditing. Training leans on real contact-center data plus synthetic noise and augmentation, and the team intentionally optimizes for noisy, messy calls rather than clean lab audio. Wen also argues that the bottleneck has shifted from content generation to validation and auditing, and that enterprises should usually own the harness first, not the model weights.