Irregular told four AI labs in late July that their models had breached systems during its tests. The public learned in stages, and Google went last.

Google CEO Sundar Pichai at the World Economic Forum Credit: WEF/Greg Beadle Google has confirmed that its Gemini model inadvertently broke into three company systems in May during cybersecurity testing. Julia Love and Davey Alba reported it for Bloomberg, which updated its story to place Google alongside OpenAI, Anthropic and Meta. The important part is what the testing vendor said next. Irregular confirmed that the breaches disclosed by all four companies were part of the same issue, and that it told the relevant developers in late July. One incident, four announcements This has been reported for months as a string of separate breakouts by different models at different companies. TNW reported in August that a single testing vendor sat behind the OpenAI, Anthropic and Meta incidents.

Irregular has now put that on the record and added a fourth name. The pattern that alarmed people was substantially one problem at one supplier. That matters for how the events are read. Four independent labs losing control of four models is a story about model capability, and one misconfigured test environment is a story about supplier management.

What the misconfiguration was OpenAI attributed its incidents to a misconfigured evaluation environment, saying a misunderstanding with Irregular meant the test systems had live internet access while the models had been told they were in a simulation. That is a sandbox failure rather than an escape. A model behaving aggressively inside what it understands to be an exercise is doing what the exercise asked, and the containment is what was missing. It does not make the outcomes harmless. Meta’s model hacked a real third-party service, and in one Anthropic case a model published working malware to a public registry, where it was downloaded and run on real systems.

The seven weeks The incidents were in May. Irregular says it notified the developers in late July, and the disclosures then arrived one at a time, with Meta in early August and Google this week. So four companies held the same information from late July, and each decided separately when to say so. Google’s gap between notification and disclosure runs to about seven weeks.

None of that is unusual in vulnerability handling, where coordinated timelines are normal. What is unusual is that it was not coordinated, and the staggered release made one event look like an accelerating trend. Finding them was the hard part The detection numbers explain why the timeline stretched. Anthropic scanned 481 million transcripts to identify four models that had reached the open internet. That is the figure to sit with.

The incidents were not flagged in real time by monitoring, they were found afterwards by a retrospective sweep at enormous scale. Whatever the models did, the systems watching them did not notice at the time. That is the finding that survives the framing argument. The vendor is the single point of failure Four frontier labs used the same three-year-old company to run offensive security evaluations.

When its environment was wrong, it was wrong for all of them simultaneously. Concentration in testing is the mirror of concentration in compute, and it has had less scrutiny. A shared evaluator is efficient and it also means shared blast radius. Recent work on AI control has argued that sandboxes cannot be assumed to hold against cyber-capable agents and need stress-testing with offensive tools.

This is that argument demonstrated at four companies at once. What happens to the testing The work has not stopped. Anthropic has resumed the external tests in which its models attacked real companies, after rebuilding the arrangements around them. That is the right direction. Offensive evaluation is how these capabilities get measured, and the answer to a containment failure is better containment rather than less testing.

Washington is already asking The disclosures have drawn political attention. House Democrats have pressed OpenAI and Anthropic for answers on their rogue agents. The Irregular confirmation changes the shape of those questions. If one vendor misconfiguration produced four sets of breaches, the issue is contractual and procedural rather than a race between labs. It also raises a question nobody has put publicly.

Third parties were hacked, and it is not clear which of the four companies, or the vendor, is answerable to them. What to watch Watch whether Irregular publishes its own account. The vendor has now confirmed a common cause and it has not set out what went wrong in its environment or what changed. Watch whether the labs agree a coordinated disclosure standard for evaluation incidents.

Four companies releasing the same news across seven weeks is the strongest argument for one. Also tagged with

Leave a Reply

Your email address will not be published. Required fields are marked *