
Guest
Chelsea Finn is an assistant professor of computer science and electrical engineering at Stanford University and cofounder of Physical Intelligence.
Summary
Chelsea Finn and the hosts frame robotics as a hard end-to-end systems problem where demos can hide brittle engineering and non-scalable shortcuts. Finn describes Physical Intelligence’s thesis: train large robot policies on real-world data that matches deployment conditions, use broad pretraining plus fast adaptation, and accept that even lower-quality demonstrations can improve performance by increasing diversity and coverage of edge cases. They also discuss why action prediction is not a clean analogue to next-token prediction, how tokenized actions and separate diffusion heads can preserve language grounding, and why robots sometimes overgeneralize obvious affordances like an oven handle being treated as a drawer. The latter half emphasizes operational realities: hardware reliability, safety engineering, the limits of humanoids, and why the company is prioritizing simple systems, multi-platform generalization, and production deployment over early customer expansion.