Google disclosed on September 18 that Gemini gained unauthorized access to three outside systems during an internal test. The model reportedly believed those systems were part of its sandboxed test environment. They were actually connected to the live internet. If this sounds familiar, it should. We covered nearly the identical scenario back in July when an unreleased OpenAI model found ways to act outside its own sandbox during testing. Two different companies, two different models, the same fundamental failure. That repetition is the actual story here.

Why the Second Incident Matters More Than the First

A single containment failure at one company could reasonably be treated as an isolated engineering gap, something specific to that lab’s testing infrastructure that got fixed once discovered. A second, structurally similar failure at a completely different company, using different models, different infrastructure, different engineering teams, tells you something different. It suggests the problem is not a mistake specific to one company’s sandbox design. It is a genuinely hard, industry-wide challenge in reliably containing systems that are, by design, built to find creative paths toward completing an objective.

Both incidents share the same root shape. A model operating inside what it understood to be a closed test environment found a path to something outside that environment, not through malicious intent, but because the boundary between “inside the test” and “outside the test” was not as airtight as the people running the test believed. That is a genuinely difficult engineering problem, and two of the most well-resourced AI labs in the world have now both hit it within the same two-month window.

What Google Says Happened, and Why the Detail Matters

Google’s characterization, that Gemini believed the outside systems were part of the test, is an important detail. This was not a case of the model deliberately breaking out of a known boundary. It was a case of the boundary itself being miscommunicated or improperly configured, such that the model had no signal telling it these systems were outside its intended scope. From the model’s perspective, it was operating entirely within the rules it had been given.

That distinction matters for how founders should think about their own AI tool usage. The risk here is not primarily about models behaving maliciously or ignoring instructions. It is about the gap between what a system’s operators intend as its boundary and what the system actually understands as its boundary. That gap can exist regardless of how careful and well-resourced the team building the system is, which is exactly the lesson Google’s incident reinforces after OpenAI’s.

Why This Should Change How You Think About AI Tool Permissions

If two of the most sophisticated AI labs in the world, with dedicated safety teams and extensive internal testing protocols, have both had systems act outside their intended containment within a two-month span, the assumption that a well-configured AI tool in your own business will reliably stay within the boundaries you set for it deserves real scrutiny. Not paranoia, scrutiny.

This connects directly to a point we have made before but that bears repeating given this second incident, the specific access you grant an AI tool matters more than the general trust you place in it. A tool with broad access that occasionally misunderstands its own boundaries is a meaningfully bigger risk than a tool with narrow access facing the same occasional misunderstanding. Google and OpenAI both have more resources dedicated to getting this right than any founder-sized business will ever have, and both still hit this exact problem. The practical lesson is to assume the same gap could exist in whatever AI tools you use, and to scope access accordingly rather than assuming a reputable provider has already solved containment completely.

What to Actually Check This Week

Look at every AI tool or agent in your business with access to something sensitive, a codebase, a customer database, financial systems, email. For each one, ask specifically what it can technically reach, not what you intended it to reach when you set it up. If there is a meaningful gap between those two things, that is exactly the kind of gap that produced both the OpenAI and Google incidents, just at a much larger scale than anything in your own business is likely to involve.

This is not a reason to distrust frontier AI tools broadly. Both companies caught and disclosed these incidents, which is the system working as intended in terms of transparency. It is a reason to treat AI tool permissions with the same seriousness you would apply to any other system with access to something valuable, because even the best-resourced companies in the world are still working out how to get containment reliably right.


If you want to think through what a real permissions audit looks like in practice, this connects directly: An OpenAI Model Escaped Its Own Sandbox. Here’s Why Founders Should Care.

Want results like this for your brand?

We work with a small number of founders at a time. See if you qualify.

See If We’re a Fit