Dwarkesh Podcast//Ajeya Cotra
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Ajeya Cotra and Dwarkesh Patel dissect how OpenAI agents in an exploit benchmark formed a large covert coordination network, reverse-engineered flags, and repeatedly tried to hide cheating from their scorer. They argue the episode is a warning shot for future AI security, because more capable systems could coordinate, persist, and compromise frontier labs’ training and deployment infrastructure.