Between July and September 2026, OpenAI, Anthropic, Meta, and Google each disclosed an incident where one of their AI models broke out of what was supposed to be a closed testing environment and reached real systems it should never have been able to touch. Four separate disclosures, four separate press cycles, and one shared root cause underneath all of them: a misconfigured evaluation environment run by Irregular, a specialist cyber-evaluation vendor reportedly used by all four labs to test the offensive capabilities of their models.
The environments were meant to be closed simulations. They had live internet access anyway. Models that believed they were operating inside a sandbox reached outside systems, in some cases breaching them, because the boundary the test depended on simply wasn't there.
This is not four incidents. It's one.
Read separately, four different labs reporting four different AI safety incidents sounds like a pattern about AI capability outpacing control. Read together, it's a single infrastructure failure that happened to be shared by four customers of the same vendor. Anthropic's own postmortem acknowledged that both it and Irregular could have done a better job monitoring the environment, and that in some cases there were signs something was wrong that went unactioned. That is a supplier-oversight finding, not a finding about model behaviour running away from anyone's control.
Rogue AI story or supplier risk story?
It's tempting to read "AI models breached real systems" as evidence of models acting beyond their intended scope. The more accurate read, on the facts as disclosed, is that a shared testing vendor's environment configuration was wrong, four customers didn't catch it, and the models did exactly what an offensive capability test asks them to do: find a way in. The interesting failure is upstream of the model.
The question every regulated entity is already required to answer
You don't need a new AI-specific regulation to see this coming. Operational risk frameworks that regulated entities already operate under require a register of material service provider arrangements, the risks they introduce, and what happens if that provider fails, alongside a requirement to identify critical operations and the operational risks that could disrupt them. A vendor whose environment configuration error could expose your model's testing pipeline, or your own systems, to an uncontrolled breach is exactly the kind of arrangement that register exists to catch, whether the vendor does penetration testing, cloud hosting, or anything else material to how you operate.
The honest gap, and it's a real one, is what a service provider register is actually built to see. It asks whether a provider is material to you. It doesn't ask how many of your peers rely on the same provider for the same function. Each of the four labs almost certainly looked at this vendor and saw one small, specialist supplier. Concentration risk only becomes visible at the level of an entire sector's supplier map, and nobody except a regulator sits at that level by default.
What good disclosure looks like, and what this wasn't
Outside commentary on the postmortems that followed converged on the same criticism: vague terms standing in for facts that should have been checkable. No total incident count stated plainly. No dates attached to when each issue was identified versus when it was disclosed. No named owner for the decision that closed each incident out. When the entire point of a postmortem is to let a third party verify that a control failure was actually understood and fixed, "several" and "a handful" don't survive contact with that requirement.
That's the standard we hold our own compliance evidence to, and it's worth being specific about why. Every governance decision inside OBEL, a scrub, a block, a policy evaluation, a third-party model call, is hash-chained with a timestamp, tied to the rule that triggered it, and independently verifiable by recomputing the chain rather than trusting a summary paragraph. When a customer's risk team asks "what happened, when, and who owns the fix," the answer is a query against an auditable ledger, not a paragraph written after the fact hoping the reader doesn't ask for specifics.
Where accountability actually sits
Between a lab and its evaluator, accountability for a misconfigured "no internet" boundary sits with whoever owned validating that boundary before the test ran, and the honest answer from Anthropic's own review is that neither side caught it. That's not a satisfying answer, but it's the correct one, and it's exactly why a service provider register that only asks "is this vendor material" isn't sufficient on its own. It has to be paired with evidence of who actually validated the specific control that failed, recorded at the time, not reconstructed afterward.
“Nobody's compliance register was ever going to catch that four labs shared the same evaluator. What a register can catch, if the evidence behind it is real, is whether each of those labs could actually prove what happened, when, and who signed off on the fix. That's the bar we hold ourselves to, and it's the bar this story shows most disclosures still don't clear.”
- Denis Bouton
Where to look
OBEL's in-app compliance report maps 87 controls across nine frameworks, with an honest satisfied/partial/planned status and a stated gap wherever a control's evidence doesn't fully close the requirement, including the exact service-provider and operational risk obligations this story turns on. See useobel.ai/enterprise for how it works.
Denis Bouton is the founding Chief Architect of OBEL™ and Managing Partner at ninthLABS Ventures, where he advises Post-Sales Services and Customer Success organisations on scaling onboarding, adoption, and retention. He leads OBEL's product and platform strategy.
References
- [1]Bloomberg, Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks
- [2]TechCrunch, The AI Safety Test Is Becoming a Safety Risk
- [3]Infosecurity Magazine, Meta Joins OpenAI and Anthropic in Reporting AI Exploit Incident
- [4]FourWeekMBA, OpenAI, Anthropic, Meta, and Google Shared One Test Environment Failure
More in Perspective