Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Original source
Artwork for Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Guests

Alexander PanfilovELLIS / MPI-IS PhD researcher

PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, focused on AI safety and LLM red-teaming.

Ilia ShumailovEx-DeepMind researcher

AI and security researcher; formerly a senior research scientist at Google DeepMind and a PhD holder from the University of Cambridge.

Summary

Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper on stealing reasoning traces from proprietary LLM APIs. The core claim is that encrypted reasoning blobs returned by frontier models can be replayed into smaller models from the same family, and in some cases across users and conversations, to recover hidden reasoning or induce jailbreak-like behavior. The guests say they tested Anthropic, OpenAI, and Google, and that this can leak sensitive user-session data such as passwords, API keys, emails, internal IPs, and medical details. They also discuss using reasoning traces for safety monitoring, but note that the legibility of reasoning may itself be fragile and model-specific. Much of the discussion is cautious about overclaiming: the paper is presented as preliminary, anecdotal evidence that needs controlled experiments, while the practical implication is a serious prompt-injection and replay risk rather than a clean distillation result.

Notes

Guests

Alexander Panfilov