Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Original source
Artwork for Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Guest

Neil MovvaCo-founder & CEO, Sail Research

Neil Movva is the co-founder and CEO of Sail Research, an AI inference platform for long-horizon agents.

Summary

Neil Movva lays out Sail’s thesis: AI is moving from interactive chat toward background agents that run for hours or days, so the winning infrastructure optimizes cost per token, not just latency. He argues that making tokens 10x cheaper creates a new product category and that the best user experience is often “no latency” because the work happens proactively before a human asks. The conversation then goes deep into the technical stack, including GPU throughput versus latency, batching tradeoffs, tensor parallelism, SRAM versus DRAM/HBM, and why KV cache is a persistent bottleneck even for very fast chips. Movva is contrarian on market structure too: he thinks open source and capability diffusion are durable, that enterprise adoption moves slowly, and that the premium for being a few months ahead may not last forever. Strategically, Sail’s edge comes from scavenging underutilized chips and power, building heterogeneous infrastructure, and pushing every layer of the stack to lower the price of intelligence.

Notes

Guests

Neil Movva

Hosts

Patrick O'Shaughnessy

Mentioned

Jensen Huang