How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Original source
Guest
Ming-Yu Liu is Vice President of Research at NVIDIA and leads NVIDIA Cosmos Lab focused on physical AI world models.
Summary
Ming-Yu Liu describes Cosmos 3 as a multimodal physical-AI system that starts with a vision-language model and then uses a diffusion-based generator tower to produce video, audio, and robot actions. He argues that world models should be treated as a collection of useful tools—forward dynamics, inverse dynamics, and policy—trained with shared constraints so they help one another. A major theme is temporal and embodiment alignment: Cosmos normalizes different signal rates across modalities and leverages human egocentric video as a source of transferable action-visual structure for robots. On the evaluation side, Liu says the near-term value of neural simulators is ranking preservation, not perfect realism, because they can help teams eliminate weaker checkpoints before costly real-world trials. The episode closes with deployment and openness details, including Super, Nano, and Edge variants, edge targets like Jetson Thor and Orin, and public models, code, and some data on Hugging Face and GitHub.