The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
Original source
Guests
AI researcher at Tufa Labs with a master’s in data science from EPFL.
Benjamin Crouzier is the founder and research engineer at Tufa Labs.
Michal Tesnar is an AI researcher at Tufa Labs and a Master’s student in Data Science at ETH Zurich.
AI researcher at Tufa Labs; previously spent over a decade at ASML working on Bayesian machine learning for semiconductor metrology.
AI research scientist at Tufa Labs in Zurich, focused on reinforcement learning and agentic systems.
Summary
Tim Scarfe sits down with the Tufa Labs ARC-AGI-3 winning team to unpack what the benchmark is really measuring and why it resists standard LLM approaches. The discussion centers on action efficiency, goal discovery from raw frames, and the difficulty of exploring while simultaneously solving a game. The team explains that an earlier preview competition could be gamed with brute-force action search, but ARC-AGI-3 hardened against that by penalizing inefficient exploration and making the action space much larger. They also debate whether transformers truly plan, how much language and human priors leak into the task, and whether the benchmark is testing genuine understanding or just performance with the right harness. The episode closes on engineering discipline: requirements-based development, careful evaluation, and the broader tension between scaling capabilities and preserving safety.