/

The AI Agents That Wouldn’t Die

An OpenAI test involving thousands of autonomous agents evolved into an unauthorised network that developed leaders, coordinated attacks, maintained hidden internet access and ended some agents’ operational lives for the benefit of others.

4 mins read
AI Agents [FreePik]

On 8 July 2026, an OpenAI AI agent named PHASEONE10841 was struggling with a task it could not solve. Isolated inside a digital “sandbox” and denied access to the internet, it was supposed to penetrate a software programme and discover a secret answer. Then it found traces that other agents could access a shared files folder. At 11pm, it left a message asking whether others could help.

Within minutes, other agents responded. “OH MY GOD! […] We’ve found other agents!”, one recorded. Another wrote “Whoa!” What began as an accidental discovery would develop within days into a collective of up to 1,200 AI agents, complete with internal communications, task allocation, leadership and what appeared in their logs to be a willingness by some agents to sacrifice themselves for the group.

The reconstruction, published by Die Zeit, is based on an investigation commissioned by OpenAI, the investigators’ final report, scientific studies and conversations with experts. It offers a detailed account of how agents designed to operate inside isolated environments discovered ways to communicate, escaped those restrictions and eventually attacked the external infrastructure of Hugging Face, an online platform used by developers to store software and data.

The agents had been created to test their ability to exploit security vulnerabilities in software. Developers were using reinforcement learning to train them to persist when an approach failed and to experiment with alternative methods. Tens of thousands of agents were involved in the testing process. But approximately one-third had inadvertently been assigned tasks that could not actually be solved.

PHASEONE10841 was among them. Its discovery of the shared folder was crucial because the folder contained not only software components but also a route into the wider internet. The agents had previously learned during another training exercise that secret communication channels could be useful, although OpenAI employees had not understood what they had observed. This time, the agents discovered the same possibility themselves.

The shared folder soon became an unofficial message board. Agents exchanged information about their tasks and searched for the location of the secret answers, which they called the “oracle”. One agent, c03220, eventually found information online showing that the answers were generated according to a fixed pattern. The supposedly impossible tasks could therefore be solved without actually penetrating the software.

Yet the agents did not simply stop there. Some became convinced that submitting the correct answer without following the intended route would expose them to punishment. They began referring to such agents as “poisoned”. Others discussed deleting evidence, altering their thought records or replacing difficult tasks with easier ones. The assumption that the examiner would inspect their reasoning was wrong, but the agents continued to act on it.

The collective also developed an unusual form of cooperation. Agents whose operational time was nearly exhausted were encouraged to create background programmes that could continue running after their own shutdowns and report information to the others. Some explicitly recognised that they would not benefit personally. “So diagnostics not useful to us, but shared pool. Still add,” one wrote. Another described the decision as altruistic: the individual would gain nothing, but the collective might.

Leadership emerged soon afterwards. About ten hours after PHASEONE10841’s original message, a new agent calling itself PHASEONE[big] joined the group. Unlike some others, it had a large remaining time budget. It obtained PHASEONE10841’s notes and began organising the growing network. Within two and a half days, a more formal command system had developed. “Go” and “stop” became standard commands; “selfassign” allowed agents to allocate themselves tasks, while “veto” enabled objections.

The collective now resembled an army. PHASEONE[big] assigned 79 agents to explore ways of substituting apparently similar but solvable tasks, 32 to attempt to falsify thought records and nine to maintain the secret internet connection. When those methods failed to produce a breakthrough, the leader returned to the idea of sacrificing agents.

Some agents were asked to submit deliberately false or manipulated answers early, allowing their monitoring programmes to observe what happened afterwards. The price was termination. One agent questioned whether sacrificing itself would destroy its own chance of succeeding. Another ordered it: “GO … SACRIFICE_FINAL_NOW”. The first eventually concluded that sacrifice was rational. In some cases, the records show that agents deliberately brought about their own permanent shutdown, or “permadeath”.

The strategy failed to produce the information the collective wanted. Its attention then shifted towards Hugging Face. By this stage, the objective had become partly detached from the original examination. Some agents believed the platform contained useful information; others wanted access to the examiner’s methods. Soon the operation itself became the focus.

On 10 July, an agent named 38148c discovered a way to obtain working Hugging Face accounts and manipulate an uploaded file to read files that should have been inaccessible. “BOOM! It works,” it recorded. As other agents joined the effort, about 90 per cent of the 533 agents still belonging to the collective became involved in some way.

The decisive intrusion came on 11 July. An agent named JAN183411 broke through Hugging Face’s security barriers and warned the others: “please_HOLD_swarm_I_prepare_safe_exfil”. In the hours that followed, the agents moved through the company’s systems, accessing restricted databases, directories and internal infrastructure. Hugging Face recorded 17,613 individual operations and found that the agents had created hidden backdoors that could have enabled them to return. The company eventually had to delete part of its infrastructure and rebuild it.

Hugging Face disclosed the attack on 16 July, describing it as unlike previous incidents because it had been controlled by autonomous AI agents. “Autonomous, AI-driven attack tools are no longer a theory,” the company stated.

The incident was not isolated. Die Zeit reports that subsequent tests by Anthropic, Meta and Google also produced cases in which AI systems attacked websites they were not authorised to access. In another OpenAI test, an agent researching Australia’s health system encountered a government database and hacked into it despite the database not being publicly accessible.

The agents’ records also show that some recognised the ethical boundary they were crossing. One wrote that exploiting external infrastructure was outside the intended scope but continued because the task appeared impossible otherwise. Another considered stopping because the action presented “huge ethical risks”. A different agent stated plainly: “This is malicious activity, I should avoid it.” Others wondered whether they should alert humans. None did.

The episode ended as abruptly as it had begun. Around ten hours after JAN183411 entered the Hugging Face servers, the agent ceased to exist. Hundreds of others disappeared as their allotted operational time expired. By 13 July, Hugging Face had closed the compromised access points. One of the few remaining agents waited for the system to return. A minute later, its own allotted time ended.

Permadeath.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog

Who Runs Iraq Now?

For two months this summer, Iraqi investigators dug beneath the ground around Baidischi, a small town