How Deep Learning Finally Cracked Messy Tables - Frank Hutter
Original source
Guest
Frank Hutter is co-founder and CEO of Prior Labs and a professor of machine learning at the University of Freiburg.
Summary
Frank Hutter makes the case that tabular data is a fundamentally different modality from text or images: heterogeneous columns, missingness, outliers, and weak cross-dataset transfer made standard deep learning unreliable for years. His answer was TabPFN, a table-native foundation model trained on synthetic data drawn from structured priors, which approximates Bayesian posterior prediction in a single forward pass and avoids per-dataset training and hyperparameter search. He also explains how TabArena was built as an open, ELO-style benchmark with careful dataset curation, and why newer versions of TabPFN changed architecture to scale from small tables toward much larger ones. The discussion broadens into how coding agents can help with feature engineering while TabPFN handles the final prediction, plus ongoing work on causal inference, relational data, time series, and the TabPFN 3.5 release.