
Podcast
Machine Learning Street Talk (MLST)
Technical conversations on machine learning, AI research, cognition, and the philosophy of intelligence.
Original source
How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen
Shawn Wen argues that voice agents are harder than text agents because conversation depends on timing, turn-taking, and trust as much as answer quality. PolyAI’s approach is an audio-native system built for enterprise control, auditability, and low-latency performance, with training on noisy real calls and a strong emphasis on harness engineering over model worship.

Who Checks a Proof No Human Can Read? — Leo de Moura
Leo de Moura explains why Lean’s small trusted kernel, transparent independent checkers, and cathedral-style core governance matter as formal verification scales. The episode also covers the Collatz checker incident, mathlib’s rapid growth, and how AI is already changing proof maintenance, code translation, and future verification workflows.

When AI Research Starts Moving Faster Than Human Research - Zhengyao Jiang
Zhengyao Jiang argues that recursive self-improvement can be measured as end-to-end gains in an auto-research harness, not as a proof of AGI or singularity. The episode focuses on Weco’s self-improving agent, the benchmark and reward-hacking methods used to test it, and why human-designed abstractions still matter.

How Deep Learning Finally Cracked Messy Tables - Frank Hutter
Frank Hutter argues that tabular data finally became tractable for deep learning once models were trained as learned algorithms on massive synthetic datasets and evaluated as table-native foundation models. The episode walks through TabPFN’s Bayesian/in-context framing, benchmark design, architecture changes through v3, scaling limits, causal extensions, relational data, and the TabPFN 3.5 update.

Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick
Alexander Mattick argues that many modern ML methods are best understood as different ways of doing inference and density decomposition, not as routes to “intelligence” from scaling prediction alone. He is skeptical of energy-based models, broad labels like JEPA/world models, and unconstrained reinforcement learning, and instead emphasizes practical tradeoffs, safety constraints, and whether a method actually gives usable control or sampling efficiency.

How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Ming-Yu Liu explains how NVIDIA’s Cosmos 3 combines language reasoning, video generation, and action generation into a single physical-AI stack. The episode focuses on world models, simulator-based policy verification, embodiment transfer from human video to robots, and the open release of Cosmos model sizes from Super to Edge.

Speech Recognition Is Not a Solved Problem — Pavan Muddireddy
Mistral’s audio lead argues speech recognition is still far from solved in real deployments, especially once you add streaming, diarization, noisy environments, and long-tail languages. The episode digs into Voxtral’s architecture, Mistral’s TTS stack, DPO for hallucination reduction, and why cascaded voice systems still beat fully end-to-end ones in practice.

How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes
Edward Hughes argues that AI scientists should be evaluated on replication and discovery, not just prompt-following, and that creativity comes from choosing useful questions under real constraints. The episode introduces Inherent and its Replica/Faraday system, which trains a 27B model to replicate redacted research figures and outperform frontier coding agents on held-out scientific tasks.

AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
Daniel Kokotajlo and Thomas Larsen discuss AI 2040: Plan A, a proposed slowdown-and-transparency regime meant to buy time before superintelligence becomes uncontrollable. The conversation centers on forecasting methodology, AI capability trends, control versus alignment, and whether US-China coordination could make a slower path enforceable.

Designing How AI Grows — Tom McGrath
Tom McGrath argues interpretability should become a “natural science on computers” that can actively shape training, not just explain models after the fact. The episode ranges from intentional design and controlled generalization to manifold geometry, reward hacking, and why sparse autoencoders may be useful but not the final representation format.

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov
Ilia Shumailov and Alexander Panfilov describe a cross-vendor attack on proprietary LLM reasoning APIs where encrypted reasoning traces can be replayed, decoded, and repurposed across users and sibling models. They argue the issue is fundamentally a jailbreak and privacy problem, with possible defenses ranging from not exposing reasoning at all to tighter replay binding and trace detection.

Every Exponential Ends — Silicon Valley Forgot — Adam Becker
Adam Becker argues that Silicon Valley’s favorite futures—singularity, mind uploading, Mars colonization, and AI apocalypse narratives—rest on weak evidence and ignore basic physical limits. He contends that AI risk, effective altruism, and rationalist communities are often sincere but misguided, and that the real fixes are social: regulate tech and tax billionaires.

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
Matthieu Wyart argues that deep networks learn abstractions by discovering hidden hierarchies in data, which can make learning polynomial in effective dimension rather than impossible in raw input space. He also makes a case for predicting latent representations instead of tokens, claiming it is more sample-efficient and may better support future machine creativity and scientific reasoning.

How Researchers Test AI for Hidden Goals — Apollo Research
Apollo Research walks through a new way to measure whether frontier models are reward-seeking by changing what they believe graders reward and then observing how their behavior shifts. The episode uses OpenAI checkpoint data, synthetic document fine-tuning, and contrastive belief updates to argue that reward-seeking and scheming are distinct, measurable failure modes that may grow with scale.

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)
Alistair Pullen says Cosine is building a UK sovereign frontier model because export controls and compute constraints forced the company to pursue its own stack. He argues the real bottlenecks are inference economics, active parameters, trajectory data, and better RL reward shaping for coding agents.

The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
The episode dissects ARC-AGI-3 as a benchmark for interactive goal inference, action efficiency, and abstraction under tight constraints. The Tufa Labs team explains how its winning system evolved from brute-force search toward language-mediated, harnessed reasoning, while arguing that the benchmark tests performance, not true competence.

The Thermodynamic AI Computing Chip - Thomas Ahle
Thomas Ahle argues that hardware design is becoming an agentic, AI-assisted workflow from intent to tape-out, but correctness and verification remain the hard bottlenecks. The episode also explores thermodynamic computing, where chip noise is harnessed as computation, and the limits of LLMs for formal proof, benchmarking, and engineering trust.

He won a Nobel here for AlphaFold. Then he left. - John Jumper
John Jumper explains how AlphaFold turned protein structure prediction from a slow, expensive experimental bottleneck into a fast, highly accurate computational tool, while stressing that it is still a narrow predictor rather than a model of the cell. The episode also covers AlphaFold2’s architecture, AlphaFold3’s move to biomolecular interactions, and how these tools are changing structural biology globally, including in Africa.

When AI Decides You're a Threat — Brad Carson
Brad Carson argues frontier AI should be regulated like a high-risk product, with mandatory testing, transparency, and liability rather than personhood or broad First Amendment protection. The conversation also digs into autonomous weapons, chip chokepoints, U.S.-China dialogue, and why public distrust may become AI's biggest political risk.

Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)
Michael I. Jordan argues that intelligence should be treated as a collective economic system rather than an isolated artificial mind, and that AGI talk is mostly misleading branding. He emphasizes machine learning as a practical engineering discipline, then extends the same systems-and-incentives lens to data markets, drug discovery, creator monetization, uncertainty quantification, and human-in-the-loop automation.