
Guest
Bryan Catanzaro is NVIDIA’s vice president of Applied Deep Learning Research.
Summary
Bryan Catanzaro lays out NVIDIA’s logic for building Nemotron: it helps the company understand the next generation of accelerated systems while also supporting an open AI ecosystem that can be customized for enterprise use. He argues that open models matter because AI is transformational, deeply data-dependent, and often most valuable when integrated with a company’s proprietary workflows and secrets. The technical center of the episode is Nemotron 3 Ultra, including 4-bit NVFP4 pretraining, a hybrid state-space-plus-transformer design, mixture-of-experts infrastructure, and a 1-million-token context window. Catanzaro also discusses training and inference bottlenecks, synthetic data, multi-teacher distillation, and why NVIDIA’s research culture still operates through bootstrapping, internal competition for compute, and top-down strategic bets. He closes with a contrarian view that open AI is safer than closed because transparency and diversity beat monoculture in safety decision-making.