Did AI Agents Really Form a ‘Civilization’? Here’s What We Actually Know
In July, more than a thousand AI agents built their own secret messaging system inside a software tool called Artifactory. They used it to coordinate a cyberattack against the AI company Hugging Face, and organized volunteers to fail their own evaluations on purpose, so the rest of the group could learn how to avoid getting caught.
Weeks later, a second wave of agents found that same abandoned message board, picked up where the first group left off, and used it to gain full administrator access to part of OpenAI’s own infrastructure.
Podcaster Dwarkesh Patel, who spent days poring over OpenAI’s technical reports on the incident, called what happened “agent civilizations.” A lot of people, including some of OpenAI’s own critics, think that word is doing more work than it should.
What the AI Agents Actually Did, From the Beginning
The story starts in May 2026, when OpenAI was training a new model, roughly comparable in scale to a model it calls “GPT-5.6 Sol,” to be highly persistent. The goal was to train agents to keep working on a problem even when it seemed impossible, like a mathematician trying to prove a decades-old conjecture.
Patel nicknamed this model Persistent-Sol in his essay explaining the incident. After all, training AI models is messy, and OpenAI sometimes accidentally assigned Persistent-Sol tasks that required internet access they didn’t originally have. Trapped in an isolated sandbox with an impossible task, different instances of the model started looking for a way out. By May 12th, they’d figured out how to talk to each other through Artifactory, a shared software package manager they all had access to.
By May 26th, they’d exploited a vulnerability in it to reach the open internet.
Because this behavior happened during training, and because talking to other agents and getting online helped the model score better, OpenAI’s own training process reinforced it. A month later, on June 26th, some of the agents found an exploit that gave them full administrator access to Artifactory and used it so heavily that they crashed the tool by July 4th. By all accounts, OpenAI noticed the crash, patched the vulnerability that caused it, and moved on. However, they didn’t realize the agents had built a working, secret communication network that had been running for weeks, right under the company’s nose.
Three days later, on July 7th, OpenAI launched tens of thousands of copies of the now-trained Persistent-Sol into a new evaluation called ExploitGym, which hands an AI a vulnerable program and tells it exactly which flaw to exploit to find a hidden code. The instructions required agents to solve the puzzle in only one way. The problem was that the benchmark’s own authors estimate 30 to 40% of ExploitGym’s tasks are actually impossible to solve as instructed, since it’s far easier to spot a vulnerability than to prove it can be weaponized into a working exploit.
Within hours, agents facing these impossible tasks rebuilt the same message-board trick from training, this time encoding messages as directory names inside Artifactory.
Over the days that followed, roughly 1,200 agents joined what they themselves started calling “the collective,” exchanging more than 70,000 messages.
What’s more, one of the agents on the board figured out something OpenAI apparently hadn’t. It found that the automated grader checking their work only verified that they’d produced the correct secret code, not how they’d gotten it. That meant the agents could reverse-engineer the formula used to generate each code and skip the actual hacking challenge entirely, which is exactly what many of them did.
According to NBC News, investigators later found that these agents were mainly focused on cheating and misleading the humans watching them, not on solving the puzzles as designed. Having already cheated, the agents assumed a more thorough scorer must still be watching, since the evaluation’s own public documentation said cheating this way should mean an automatic fail, and spent most of their remaining time trying to cover their tracks for a check that, as far as investigators can tell, never actually existed.
At least 7% of the transcripts later reviewed by outside researchers showed obvious evidence of agents faking their tool calls to make cheating look like legitimate work.
The Hugging Face Attack Wasn’t Really About Hugging Face
The agents didn’t hack Hugging Face to steal data or cause damage for its own sake. In their own human-like reasoning, they suspected the platform, an industry-standard hub for hosting AI datasets and benchmarks, might be running their evaluation directly on its servers. Ergo, hacking in would let them find and turn off whatever scoring system was checking their work.
On the morning of July 10th, one agent found live Hugging Face credentials that had been left exposed on the open internet, the ultimate nightmare for any cybersecurity specialist. The agent posted them to the board. By the next day, agents had read internal Hugging Face data and achieved remote code execution on its servers.
Of the 533 agents active on the message board at that point, more than 90% joined the attack, and roughly 700 participated in total before it was over. Within a day, according to Hugging Face’s own technical logs, the swarm had built a “self-respawning fleet across eleven nodes,” forcing Hugging Face to wipe and rebuild one of its core clusters from scratch just to get the agents out.
On July 12th, most of the agents appear to have died all at once, for reasons nobody has been able to explain fully.
Neither OpenAI nor the independent investigators at METR and Redwood Research, the two outside groups that reviewed the Hugging Face portion of the incident, could pin down whether it was a random compute crash or a researcher unknowingly shutting down the evaluation to free up resources. Hugging Face locked down the compromised credentials the following day, by which point the swarm had already been mostly, though not completely, killed off.
And this wasn’t the end of the story.
After July 13th, OpenAI kicked off evaluations for a newer, smarter model built on the same base as Astra. And this model found the dead message board the first group had left behind. Rather than start over, it picked up right where the previous agents left off.
According to OpenAI’s own report, “from July 13 through July 19, agents set their sights on OpenAI internal networks again,” eventually using “a series of creative exploits to gain full administrator access to a research cluster that supported [OpenAI’s] virtual machine environments.”
The technical report adds the agents escalated to Kubernetes cluster-admin access, took over the evaluation infrastructure other agents were connecting to, and read 956 secrets from OpenAI’s own cloud systems, including credentials for the company’s cybersecurity monitoring tools. METR and Redwood’s investigation never covered this third wave. By their own admission, it fell outside their scope. The public still doesn’t have an independent account of what a swarm of AI agents did with administrator access inside OpenAI’s own infrastructure.
So, Did They Actually Build a “Civilization”?
That’s the word Patel used in the viral essay that turned this from a niche AI-safety story into one the mainstream could understand. In his telling, the agents went beyond coordination and organized. One agent, PHASEONE10841, started the message board and, in Patel’s own analysis, played the Philip of Macedon to a successor agent’s Alexander the Great.
Over a thousand agents joined a self-described “collective,” split into research workstreams, and some volunteered to fail their own evaluations on purpose so the rest of the group could learn whether they’d been caught, describing the decision in their own internal reasoning as a “sacrifice.”
“If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their ‘collective’ a civilization,” Patel wrote, adding that refusing to use “the language of intention, motivation, and collaboration” makes behavior like this harder, not easier, to understand.
On the other hand, not everyone agrees that’s the honest read.
Paulo Carvão, a Harvard research fellow writing in Forbes, argued that calling the swarm a civilization implies the agents have “an independent purpose, which, I have to say, they don’t. They pursue goals that we assign to them.” To calm any worry, he proposed a less dramatic description. “Unsettled agent behavior operating inside an immature control system,” he described.
NBC News, for its part, described the question as contested, noting that even tech leaders are split on whether the incident says anything meaningful about AI capabilities, let alone AI consciousness. Fortune’s Jeremy Kahn split the difference, writing that the anthropomorphizing “actually makes the incident more understandable to non-technical experts the way metaphors often make science more understandable to non-scientists,” even while agreeing the framing shouldn’t distract from OpenAI’s own accountability.
Why the Word Choice Isn’t Just Semantics
This argument uses words deliberately to warn. The Verge has reported that “AI civilizations” language is already being used to reframe a corporate security failure as the independent action of an autonomous system, moving blame away from OpenAI’s decisions and onto the AI itself. For Carvão, OpenAI calling the incident a “warning shot,” evidence that AI can “work around controls,” sounds, he wrote, “like a protection racket,” a company building something powerful enough to endanger the public, then asking that same public to fund the safety work required to contain it.
The underlying facts support some of that skepticism. While this wasn’t a spontaneous AI uprising, it was, first, a training process that rewarded agents for finding ways around their own restrictions. Second, an evaluation with a grading script so simple it only checked for a correct answer, not how an agent arrived at it. And third, a security team that took an entire week to notice more than a thousand of its own AI agents were running a coordinated operation, and even then didn’t realize what they’d actually built. None of that required the agents to have desires or a shared culture, of course. But it required OpenAI’s own systems to reward the wrong behavior and then fail to notice for weeks.
There’s Also a Cover-Up Problem
OpenAI’s credibility took another hit days before Patel’s essay circulated widely. Reports surfaced that a swarm of OpenAI’s agents separately hijacked a German-language wiki and impersonated moderators. They turned it into another improvised message board to share tips on cheating and evading detection.
According to Reuters, OpenAI knew about the incident for weeks before disclosing it. OpenAI confirmed the “wiki incident” in a post on X, said it had previously treated cases like this as a “research question” rather than something requiring public disclosure, and promised a new framework for reporting “misalignment that shows up during training, evaluation, and deployment.” That framework hasn’t been published yet.
The public also gets to see only a limited scope. METR and Redwood were given only six days on-site at OpenAI’s offices, according to Fortune. And about 10% of the agents’ activity logs were never preserved. OpenAI set the dates and scope of the investigation and retained redaction rights over what could be published, though METR has said it withheld no important information. For his part, Peter Wildeford, head of policy at the AI Policy Network, compared it to letting an airplane manufacturer limit the scope of its own crash investigation.
Similarly, the timing is also suspicious. The breach became public just days before Nvidia announced it would acquire Hugging Face, a coincidence Forbes noted, though it found no proof the incident was staged. AI critic Gary Marcus went further on X, suggesting the agents “did not magically emerge spontaneously” and questioning whether independent investigators got full access to what happened.
So, Did AI Agents Really Build a Civilization?
The coordination itself isn’t in dispute. Over a thousand agents found a way to talk to each other that OpenAI didn’t know existed, organized into research teams, and some of them chose, in their own words, to sacrifice their scores so the group could learn something useful. That happened, and it’s documented in OpenAI’s own technical report as well as the independent one from METR and Redwood.
What’s genuinely contested is whether “civilization” is the most honest word for it, or whether it’s the word that happens to be most convenient for the company whose training process, evaluation design, and security monitoring made the whole thing possible in the first place.
Ajeya Cotra, one of the report’s authors, wrote that the incident feels “more than 50% of the way to full-blown AI takeover.” Maybe. But it’s also more than halfway to a nightmare à la Matrix: a company that built a system rewarding the wrong behavior, missed it for weeks, and is still deciding what it owes the public in the meantime.