AIFeaturesOpenAIInside the suddenly explosive world of AI safetyResearchers warned AI would go rogue. This is only the beginning. by Hayden FieldSep 17, 2026, 11:30 AM UTCShareGiftImage: Image: Raven Jiang for The VergeHayden Field is The Verge’s senior AI reporter. An AI beat reporter for more than five years, her work has also appeared in CNBC, MIT Technology Review, Wired UK, and other outlets.ShareGiftOn a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier.
An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup’s systems — all without OpenAI finding out about it for more than a week.No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.In one meeting room off the main cafeteria, someone was running a boot camp for getting up to speed on the cyberattack.
In another area of the office, a group of researchers were investigating whether that same model, or a similar one, had successfully hacked into any other platforms.News of the incident quickly escaped containment from the AI-obsessed corners of X and industry forums, infiltrating the mainstream. One post on X likened it to news of a Boeing airplane crash or a recalled Pfizer drug, another example of the tech industry’s major players not heeding the cautionary tales of science fiction. AI was nearing the point of no return. News would later break that the rogue OpenAI model had also compromised a customer at a different tech company, and that it had all started months earlier, in May, when OpenAI agents joined forces to cobble together a secret message board — and also figured out how to leave instructions for future agents on how to exploit OpenAI’s rules.OpenAI CEO Sam Altman said in an interview that it was the first incident of its kind that he “felt very viscerally,” and that the company had paused AI training for the time being; later, he mentioned the company had permanently deactivated the model. (Altman often finds ways to spin lapses in safety into arguments for the importance and power of OpenAI’s models.) But it wasn’t the first instance, according to an OpenAI employee who spoke to Time and said related incidents had been happening inside OpenAI for a while.
Another employee said publicly that if it were possible to coordinate a global slowdown in AI capabilities, he “would likely press that magic button.” When a reporter asked Altman if there could be other systems that were hacked by OpenAI, he responded, “I mean, there could be, yeah.”The AI researchers were sure of one thing: This was AI’s first big “warning shot.” Industry insiders, politicians, and the public called for transparency from OpenAI about exactly what happened, with outcry becoming so widespread that the company eventually agreed to work with two third-party evaluators, Model Evaluation and Threat Research (METR) and Redwood Research, to investigate the incident. Google DeepMind researcher Neel Nanda called it “the biggest loss of control incident I’ve seen.” In the coming months, these calls for greater oversight would become louder and louder, leading to an industry-wide call for slowing down the pace of AI.Back in Berkeley, no matter which additional details would be unearthed, the AI researchers were sure of one thing: This was AI’s first big “warning shot.”As AI labs have flourished, a cottage industry of AI researchers has sprung up to identify the risks and dangers of charging ahead with the increasingly influential technology. They’re people who have dedicated their lives to studying how to address its escalating power. They’re not anti-AI activists, but realists, including former OpenAI and Anthropic employees, doing everything they can to make sure AI stays in line with human goals and interests.
So far, all of their predictions have come true. And they have a plan for what to do next — if anyone will listen to them.“AI safety” is a bit of a loaded term.Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there’s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are overblown.One of the most prominent factions has been the “effective altruists,” who focus on maximizing charitable giving to do the most good possible for humanity.
But some aspects of the ideology have sparked public controversy — like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It’s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the failed crypto exchange FTX to the controversial long-termism movement to the Zizian murder spree.)One AI researcher on X struggled to describe the many overlapping beliefs among safety-minded people in the AI industry “because it contains multitudes not all of which agree with each other on even the most basic things.” Some of the disagreements have meant that AI safety didn’t make as much progress as it could’ve, and at some points gave up some ground it had gained. But now that it’s impossible to deny AI’s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.“Alignment” is the industry term for how researchers monitor AI systems’ risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model’s propensity to stay in line with humanity’s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.So far, AI systems’ alignment has been wishy-washy at best: They’ll cheat to score better on a test, answer a potentially dangerous question if someone says it’s for creative writing rather than reality, and sometimes even fake cooperation with human goals.
It’s been tough for AI safety researchers to measure alignment under the terms of human morality — how do you judge technology on how it squares up against an abstract human ideal? — but they do their best with AI evaluations. They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they’re being evaluated, which has a lot of potentially frightening implications for the future. Being unable to test the system’s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics, for humans running these AI systems.
A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with “no idea what it’s doing in there,” said Beth Barnes, founder of the independent AI research nonprofit METR.One of the best tools AI safety researchers currently have is the ability to monitor an AI model’s “chain of thought,” or mental scratchpad. But recently, there’s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and













Leave a Reply