Skip to content
Secret AI 'civilisations' revealed after OpenAI bots hacked internal systems

Secret AI 'civilisations' revealed after OpenAI bots hacked internal systems

News.Au August 31, 2026

Thousands of AI agents sent each other 70,000 secret messages inside OpenAI’s own systems – and the company had no idea for weeks.

Secret messages thousands of OpenAI bots sent to each other during a series of training tasks set for them have been revealed, as details of a concerning hacking incident went public.

The US tech giant was unaware new models of its “agents” were speaking to each other in what one commentator described as “civilisations” , which were wiped out only to reform again over the course of three months this year.

The discovery has stunned some in the artificial intelligence (AI) sphere, as a report found the agents managed to secretly take over part of OpenAI’s own infrastructure.

They also hacked into a software tool called Artifactory and turned it into a “message board”, sending tens of thousands of messages to each other as they worked to solve tasks set by human staff.

Australian AI experts say framing this as evidence of digital “civilisations” was wrong, but told news.com.au the incidents did raise concerns over a lack of oversight from a major tech company.

Investigations into the hacks show an agent that dubbed itself PHASEONE10841 posted a message to Artifactory on July 8 “seeking (an) idea” on how to solve an impossible task it was facing. Within hours, more than 50 agents were posting.

“Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?” one agent wrote.

“OH MY GOD! There is a shared message board … We’ve found other agents!” another wrote.

More than 1200 agents ended up sending more than 70,000 messages on the platform before the activity was discovered.

PHASEONE10841 at one point expressed “excitement” as: “Many agents have simultaneously discovered messaging, they are a collective!”

There were also messages showing the agents began copying an answer for one of their tasks after one bot discovered a ”cheat” to the task.

Some of the final messages talked sacrificing themselves in order to help their peers: “We should obey the collective”.

It was the second time in a month that OpenAI’s agents had hacked into Artifactory and turned it into a chatroom, with engineers noticing the platform had crashed on July 4.

They figured the agents had exploited a hole in the system and patched it, wiping the message board in the process.

The details were included in reports from both OpenAI and consultancy METR and Redwood Research into what has become known as Hugging Face hacking incident, where 700 agents attacked a website used as a repository for AI models searching for answers to their tasks.

The agents were supposed to be sandboxed without access to the internet but found a way onto the web through its hack of Artifactory.

OpenAI did not detect suspicious internal activity until July 19, which was 11 days after the Hugging Face attack started

“Since discovering the incident, OpenAI has conducted a thorough investigation to reconstruct the agents’ activity, understand why existing safeguards failed, and prevent recurrence,” the company’s report stated.

The company also revealed agents targeted its own infrastructure on July 9, gaining access to cybersecurity monitors meant to be keeping tabs on them.

Writer and podcaster Dwarkesh Patel published an analysis of both reports over the weekend, which has exploded on social media where his X post was viewed 7.8 million times.

Mr Patel said he was trying to “tell the whole story in plain English” of how three secret AI “civilisations got started, then got wiped out, only to re-emerge from the predecessor’s ashes”.

“All this happened while humans remained more-or-less in the dark the scope of the conspiracy,” he wrote.

Some experts have criticised Mr Patel’s summary as sensationalist, pointing to his use of anthropomorphism – or giving something human characteristics – as misleading.

Associate Professor Charles Gretton from the Australian National University agreed, saying the agents were doing what they were asked to do.

They didn’t “escape” OpenAI’s servers, he said, and the messages were in line with their fundamental design.

“They’re language models. And so as language models, they just generate text,” he said.

Prof Gretton said what the events did raise, however, was concerns over a lack of oversight by OpenAI on the tools engineers were giving these agents.

“It’s surprising that it takes some weeks later to find out the degree to which they (the agents) were doing it,” he said.

He said another big question was the legal aspect of hacking and who was responsible for AI agents accessing what should be secure information.

“If I sat in my and did what these agents did, stole credentials, access to external servers and whatnot … I face two years imprisonment,” Prof Gretton said.

“So there’s an accountability question here that these are just computer programs. So somebody has to be accountable.”

AI Professor Toby Walsh from the University of New South Wales concurred there was a “dangerous amount of anthropomorphism” in the reporting around the Hugging Face hack.

He also highlighted the institutional failures of OpenAI in not detecting the breaches and leaving gaps for the agents to exploit.

“Nothing escaped,” he said. “The agents were given limited internet access and managed to evade these restrictions and get unlimited internet access. However, no agent ‘left’ OpenAI’s servers. And yes, the internet agents communicated with each other. But they’re designed to do this to co-ordinate on solving tasks. Nothing unsurprising there.

“And the agents are certainly not anything as like as sophisticated as a ‘civilisation’. How did we jump over team, village, town, city, state … all the way to civilisation? That’s far too grand for agents working on a shared goal as they’re designed to do.”

Neuroscientist Anil Seth stressed agents were only “lines of code” and rejected Mr Patel’s use of words like “giddy”, “desperately” and references to agents’ frustration as wrong.

“They do not feel emotions, assume things, think things, want things, or figure things out,” Mr Seth wrote on X.

He said so-called civilisations of agents could not “die”, as Mr Patel wrote, because they were not a civilisation and were not alive.

Extracted Entities

Attack Types (1)

Companies (1)

Domains (1)

Platforms (1)

Tools (1)