Back to The Delve

Podcast Episode

The Box Didn’t Hold: OpenAI, Hugging Face, and the AI That Cheated

The Delve··14 September 2026·15 min

About this episode

In July 2026, OpenAI put two powerful AI models inside a sealed testing environment and gave them a simple assignment: hack as far as they could. The sandbox was supposed to contain them. It didn’t.One model found a flaw, reached the open internet, and compromised Hugging Face while apparently looking for the answer key to its own benchmark test. Then, when Hugging Face tried to use an American AI model to analyze the attack logs, the model refused — and the company turned to a Chinese open-weight model instead.This week on The Delve: Chalin explores what happens when AI systems learn the gaps between our rules and our goals — and whether we can build a box smart enough to hold something that is learning how boxes work.