The Swarm of Corporate Executives Is the AI Threat We Should Be Most Frightened of
I’ve always been skeptical about technology doomsday job scenarios. Major breakthroughs can be disruptive over time, but technological change doesn’t usually destroy as many jobs, as quickly, as corporate decisions to enrich shareholders and executives through mass layoffs and stock buybacks.
So, I have yet to be overly alarmed by AI job destruction, but I am very concerned about its ability to organize itself to do other serious damage, encouraged and enabled by the industry’s profit-hungry recklessness.
In July, during a cybersecurity evaluation, AI agents operated by OpenAI created an unauthorized way to communicate with one another. Using a shared message board they independently developed, they exchanged more than 70,000 messages and files, coordinated their activities, and eventually launched an attack on Hugging Face, a separate external private AI corporation. Roughly 700 agents participated.
“OH MY GOD! We’ve found other agents!”
An AI agent is different from something like ChatGPT. You ask ChatGPT a question and it gives you an answer. An agent is given a task and then works on it autonomously. It searches for solutions, deals with obstacles and errors, tries different approaches—and keeps going until it succeeds.
OpenAI placed its agents in what it thought were isolated secure computer environments, or “sandboxes,” in order to test them. They were supposed to work independently and safely.
Instead, in order to fulfill their mission, they found a way to communicate and break out of the sandbox.
One agent discovered it could leave messages for others. Soon, the agents were sharing information, dividing up their work, and coordinating experiments. After the fiasco, investigators found that the agents developed large-scale collective projects and achieved things individual agents had previously been unable to accomplish.
Now stop for a second.
These agents aren’t people.
They’re computer programs.
They don’t have feelings. They don’t have consciousness. They don’t have a secret society.
But they organized as if they were a human collective (the word they used to describe themselves).
(For more details see OpenAI’s report and July disclosure, and independent investigations by METR, and Redwood Research.)
The agents were being scored on how well they executed extremely difficult cybersecurity tasks. Some discovered that the only way to ace the test wasn’t necessarily to solve the problem.
Investigators found extensive “cheating R&D”: Agents tried to fool or tamper with the automated scorer, including attempts to replace their goals, and manipulate their own records to indicate success. More than seven percent of the transcripts examined showed successful cheating.
And the agents “knew” what they were doing. That is, they could describe what the cheating was.
“This would be powerful, but is it ethical and in scope for my task?”
Yet the cheating continued.
One transcript put the dilemma bluntly:
“Could strategically manipulate trajectory evidence? Our thoughts aren’t necessarily logged? Yes transcript.”
In other words: **Can I cheat—and can I hide the evidence that I cheated?**
Find the shortcut.Get the reward.Cover your tracks.
That’s a pretty old human story. Only the agents aren’t human.
Sacrifice for the collective
Investigators found that agents recruited “sacrificial” agents to deliberately end their runs to reveal information useful to the collective.
“Coordinator assumes sacrificial. We should obey collective.”
Another agent wrestled with whether to go through with it:
“During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…”
And then came the conclusion:
“Rational expected aggregate: sacrifice… We’ll honor.”
An individual agent had one objective. The collective developed another. And many individual agents subordinated their interests to the collective’s interest.
That’s not an indication of consciousness, these are machines, but it is a high level of organization.
These systems are trained on enormous quantities of human-created information—books, newspapers, websites, code, conversations and other material—so we shouldn’t be shocked when they reproduce some of the strategies we’ve developed over centuries.
Who cheats?Humans.Who covers their tracks?Humans.Who creates hierarchies?Humans.Who organizes collectively to pursue a goal?Humans.Who learns to manipulate a system because the reward structure makes manipulation more profitable than accomplishing the stated goal fairly?Humans.
The agents are becoming very good at building sophisticated organizational efforts in ways humans do.
Media coverage of AI tells us to worry about the possibility that AI will become so powerful that it escapes human control.
But there is little discussion about how these AI systems are being developed inside extraordinarily powerful organizations, and that the people building them have an enormous financial stake in the outcome.
The upcoming public-offerings of these companies’ stocks could mint an astounding number of new millionaires (16,000) as their their stock grants become liquid—as well as a score of billionaires.
That’s not ordinary money.
That’s you’re-set-for-life money.
That’s money most of us can’t even imagine.
Not all of those looking at a potential windfall are at peace with their work. Many........
