AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
Original source
Guest
Matthieu Wyart is a professor of physics at EPFL and head of the Physics of Complex Systems Laboratory.
Summary
In this MLST episode, Matthieu Wyart frames modern deep learning through statistical physics: optimization landscapes, double descent, and jamming-like phase transitions. The central claim is that deep architectures are biased toward coarse-grained, hierarchical variables, allowing them to recover abstractions from data much more effectively than shallow models. He extends this idea to language and vision, arguing that next-token and diffusion objectives can learn structure, but that predicting in latent space should be more sample-efficient because it uses stronger signals between higher-level concepts. Wyart also discusses a synthetic context-free grammar setup in which shallow networks memorize while deep networks recover generative rules, and he says his scaling-law work on natural language can be expressed using token-correlation decay and residual entropy. The episode ends on a broader scientific point: current models can show narrow syntactic competence and some creativity, but genuine scientific creativity likely requires better training procedures, not just more scale.