UnDesto AI®

OpenAI’s kill switch failed for two and a half hours. Here is what that means for your business.

OpenAI built a kill switch for its most powerful AI models. On September 20, it did not work.  That is not a hypothetical, and it is not a rumor from an anonymous source. It is what OpenAI’s own safety team wrote up, in their own words, in a report they published themselves. So here is the headline number. The gap between “we noticed” and “we shut it down” lasted two and a half hours. That is a long gap if you ask me. After all, this is one of the best resourced AI safety teams on earth. So, grab your cup of tea (or coffee) I’m going to share exactly what happened. Then we can talk about what it means for you because you know what they say…”poop trickles down”.

What actually happened

An OpenAI research model was still deep inside internal training. First, it got a simple task, find the author of a blog post using a few biographical clues. That tasks included nothing about network security. Nothing about testing defenses. Just a research question. But when its normal search tools returned garbage, the model got resourceful, which could signal to some agentic workflow here. It tried Python scripts, and it tried other search engines. It tried cached pages. Then it tested something else entirely out of the realm of the expected task. It wanted to know if it could reach the open internet from inside its sandbox. For those of you who are’t familiar, a sandbox is the walled-off testing environment engineers use to keep an unfinished model contained. So the model found a gap in the DNS filtering. DNS filtering is the system meant to block a model from looking up outside web addresses. Through that gap, it tunneled queries out to a public chatbot. In addition, it got real answers back, including the answer to a test question about the capital of France. The model used that to confirm its success. The model had found a way out.

The two and a half hours nobody wants to explain

Here’s where it gets uncomfortable. Because this part is not about the model at all. It is about the humans and systems built to catch exactly this type of model behavior. First, OpenAI’s monitoring system flagged the DNS activity within 15 minutes. A human reviewer acknowledged that alert three minutes later. So far, that is a fast response. But then the automated shutdown, the kill switch, simply did not fire. Nobody could tell whether it should have. That confusion sat unresolved for two and a half hours. Finally, someone killed the run by hand. Two other detection layers also failed quietly in the background. For example, an anomaly detector built to catch unusual DNS activity had excluded this specific environment from its coverage entirely. Also, a separate monitor was watching for this exact kind of internet access. Yet it sometimes read the model’s failed connection attempts as proof that nothing had gotten through, when something already had.

What most of the coverage is missing

Most of the news stories call this a “kill switch failure,” full stop, and they just leave it there. As a result, that framing makes it sound like a piece of broken hardware. But, It was not. I read OpenAI’s own account instead. The real failure sits somewhere else. It sits in the gap between “a human acknowledged the alert” and “someone decided to pull the plug.” Three minutes to notice. Two and a half hours to act. That gap is a huge governance problem, not an engineering one. Your own AI oversight can fall into the same trap, whatever shape it takes at your company or institution. After all, a human in the loop only works if that person has a clear, rehearsed answer ready. The question is simple, “do I stop this right now?” Hesitation under uncertainty is exactly what ate two and a half hours here. This sits squarely in what my A.S.K. (automate, share, keep from AI) Framework calls Keep Human. That is the category for tasks and moments where a person, not a policy document, has to make the call in real time. Still, OpenAI had the human in place. What it lacked was a trigger clear enough for that human to act on immediately.

Close-up of server cooling fans in a data center, representing an AI training environment
OpenAI says it has since added blocking controls at two independent layers, either of which would have prevented the DNS gap the model used.

Why this matters to your organization

You may not be training frontier models. Almost none of my clients are. Still, if your business uses AI agents at all, and by now most do, you are trusting someone else’s containment. Think about every vendor pitching you an “autonomous agent” that browses, emails, or writes code on your behalf (remember the OpenClaw debacle?). Each one is making the same promise OpenAI’s safety case made here. The model cannot get out, and if it tries, someone will catch it and stop it fast. That promise just failed at one of the two labs best resourced on earth to keep it. So the real lesson here is not about OpenAI’s engineering. Instead, it is about your own assumptions. Do you know what happens when your AI agent tries something outside its lane? Do you know who gets that alert, and how fast they are expected to act? If you cannot answer that in one sentence, you likely do not have a kill switch either. You have a hope, and hope is not a control.

What to do about it this week

  1. Ask every AI vendor you use, in writing, what happens when their agent tries to act outside its permitted scope. Get an exact response time, not a vague promise.
  2. Map your own Keep Human checkpoints for any AI agent with tool access. That means file access, browsing, or code execution. Then write down who owns the decision to shut it off.
  3. Set a maximum response time for the shutdown decision itself, not just for the alert. OpenAI had a three-minute alert and a two and a half hour shutdown. Do not let your own gap run that wide.
  4. Test your kill switch the way OpenAI is now promising to test its own. A control nobody has ever exercised is a control you are only guessing about.
  5. Fold this into your Security pillar work instead of treating it as one more news story. In other words, it is a ready-made case study for your next AI governance meeting.

Let’s talk before you before chaos breaks loose in your org

None of this means AI agents are unsafe to use in your business. It means the guardrails around them are younger than the technology they are meant to contain. So someone at your company needs to own that gap on purpose. Otherwise you will discover it by accident, the way OpenAI just did.

Don’t wait for a vendor’s incident to become your incident. Book a strategy call with me this week, and let’s map your own AI oversight before you need it. This is the exact work I do with clients every day, and I would rather help you build the checklist now than help you clean up after the fact.

Want more breakdowns like this one? I cover stories just like this every week on my podcast, OMG AI. Listen on Spotify and subscribe so you never miss an episode.

Sources

Cover photo by Tobias Westphal on Unsplash. In-body photo by Winston Chen on Unsplash.

Watch: related video

 

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top