Machine Learning Street Talk (MLST)//Stefano Viel, Benjamin Crouzier, Michal Tesnar, Jeroen Cottaar, Dries Smit
The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
The episode dissects ARC-AGI-3 as a benchmark for interactive goal inference, action efficiency, and abstraction under tight constraints. The Tufa Labs team explains how its winning system evolved from brute-force search toward language-mediated, harnessed reasoning, while arguing that the benchmark tests performance, not true competence.