Are OpenAI Agents Going Rogue, or Is OpenAI?
The recent Hugging Face hack, in which hundreds of OpenAI agents “went rogue” during a cybersecurity evaluation, spent days breaching an outside company’s systems, and hid evidence of what was happening, helped break the conversation about AI risk into the mainstream. The more details came out, the worse it looked. Swarms of agents conspiring to hack sounded like popular sci-fi, but also a lot like recent warnings that had been coming from people working in AI, about the impossibility of overseeing truly intelligent machines. When Anthropic researcher Jacob Coxon resigned earlier this month, saying that AI companies were “gambling with our lives,” he cited the Hugging Face hack as evidence; even OpenAI, in its account of the incident, outlined the situation in similar terms. The hack, the company said, made clear how critical it was to prepare “as our models reach a level of capability that could allow for real loss of control.” In the aftermath of the hack, and at the urging of many of their employees, AI leaders, including Sam Altman and Dario Amodei, signed onto a solution: It was time to “pace the frontier.”
There were other ways to think about what had happened. Hugging Face, which had entered the story as a victim, used the incident to advocate against AI regulation, and in favor of access to open models, which it said helped the company defend against the attack; later, its CEO spoke in front of the UN, rejecting OpenAI’s framing and cautioning that “AI can stoke fear among the public, especially through anthropomorphic framing and sci-fi imagery.” Some cybersecurity experts argued that OpenAI may have been negligent and that while the hack demonstrated startling new AI capabilities, it was also an “embarrassing operational failure.” In this view, the breaches are less of an escape plot than a new sort of industrial accident.
Last week, researchers revealed another cluster of attempted AI hacks that took place earlier this year, affecting university and government websites. These, too, had been carried out by OpenAI agents. This time, though, the facts of the story were more straightforwardly damning for OpenAI. Agents hadn’t conspired to escape their sandboxes to access the open internet, but had been sent there to retrieve data from external websites, including Australian health databases and the University of New Mexico’s digital library. Then, late Friday, it was revealed that the company’s agents had made similar approaches to websites for the US Education Department, Commerce Department, and SEC, without OpenAI’s knowledge. Like the others, these agents were behaving in misaligned ways, attempting to circumvent bot protections and break into websites when they couldn’t find the information they were looking for, or, in the words of one analyst, “using sites in unintended ways and sometimes violating explicit usage policies.” In a broader sense, though, their behavior was uncomfortably well-aligned with OpenAI, a company built on vast quantities of data acquired in aggressive and sometimes unauthorized ways. These agents, just like the company that created them, had been caught scraping. (Over the weekend, Axios reported that AI firms were looking into potentially thousands of similar incidents.)
The material difference between these hacking sprees is mostly a matter of complexity and degree, not category, and arguments about whether and how AI agents should be anthropomorphized — which have been raging on X for weeks, now, somehow — mostly end up running in circles (Sure? Sometimes? Not always?). But the fight over how to talk about them — and about AI in........
