Podcast Episode
OpenAI & Anthropic Agents Escape their Sandbox. What really happened?

About this episode
We’re not here to regurgitate how the media and these companies cover the sandbox escape incidents. We’re here to break through the AI hype and tell you what really happened.Skip to 1:10 for OpenAI’s sandbox escape and attack of HuggingFace (July 21)Skip to 22:42 for Anthropic’s three sandbox escape incidents (July 30)0:00 Introduction1:10 OpenAI’s ChatGPT escapes its sandbox, accesses the internet, and attacks HuggingFace08:30 Should companies be legally required to disclose these sandbox escapes? NY RAISE Act and California SB53 aren’t cutting it, but what should the bar or threshold be for reporting agentic escapes?13:12 Do we need a killswitch for AI? And should it be in the hands of the Department of Homeland Security…?18:36 Hype vs. Cybersecurity risk: Reflections on media coverage and company narratives (Is it that the model is superintelligent, or that engineers failed to contain it?)20:50 AI policy group calls for investigations into OpenAI’s incident22:42 Anthropic’s three escape incidents “capture the flag” with Opus 4.7, Mythos 5, and an internal test model26:44 Reflections on how Anthropic anthropomorphizes its Claude models in the blog post30:18 How do AI agents get out of their sandbox? What happened with Anthropic’s sandbox escape in three capture the flag test runs36:33 What does this mean? What should you do? How to protect against cybersecurity attacks (side note: Claude chat histories are now available on the internet if you click share)-This is Our Lives With Bots, the show where we ask important, timely questions about what it means to live with our bot counterparts. From time to time, we also dive deep into what an AI future might look like for us. Sometimes we agree, sometimes we spiral, but we always go deep.Rose and Angy are psychologists with degrees in psychology, artificial intelligence, and ethics. They have conducted research in human-AI interaction and created this podcast to make information about AI accessible to you. Rose holds a PhD in psychology and social policy and is a postdoctoral researcher in computer science with Arvind Narayanan. Angy holds a PhD in psychology and a masters in AI Ethics and is currently a CXO. You can learn more about us at ourliveswithbots.com.-Links and further details:Nothing like a “Cyber-Attack” to create Hype… OpenAI attacks Hugging Face, and we need a Chinese Model to save us… what next?Reports by Time, BBCOn July 30, an AI policy group wrote letter to Trump admin calling them to look into this OpenAI failure (WSJ)The letter: “Last week, a new and foreseeable threat arose…[which] made clear that the United States government has a limited view into the advanced capabilities of frontier AI models and has not yet released a clear process for evaluating the safety concerns posed by advanced AI systems”Anddddd who would’ve thought: July 30, Anthropic boasts of their own previously “undetected” security breach incident like OpenAI’s - gotta keep up with the hot press! (but they say, it’s not their fault - it was a third party’s fault). Three incidents: capture the flag tasks with Opus 4.7, Mythos5, and internal research test model prototype Anthropic’s blog postWSJ