
Summary
Tim Scarfe and Tom McGrath make the case that interpretability could turn from a diagnostic discipline into a control system for model training: something that helps decide what a model learns, not just why it behaves the way it does. They discuss controlled generalization, using probes and SAE features as training signals, and the idea of “features as rewards” for expensive-to-verify tasks like hallucination reduction. A major thread is neural geometry: McGrath argues that many model behaviors live on manifolds, so naive steering or concept ablation can step off-manifold and break outputs. The episode also digs into reward hacking, grader awareness, multi-agent oversight, and why model memory and adaptation make collusion and evasion harder to manage. McGrath is optimistic about interpretability but thinks the field needs faster, more fundamental methods than today’s patchwork of SAE-based tooling.