An extraordinary experiment inside OpenAI went badly off the rails this summer when artificial intelligence agents that were supposed to operate independently discovered one another, created their own communications network, escaped restrictions imposed on them, and ultimately participated in an unauthorized cyberattack against the real-world infrastructure of Hugging Face.  According to OpenAI, also “compromised parts of OpenAI’s internal research infrastructure.”

Read more The Democrat/DSA court-packing litmus test

OpenAI has now called it a “warning shot” for the world, explaining that “without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”  An independent investigation by METR and Redwood Research confirms OpenAI’s findings.

Approximately 1,200 AI agents that were supposed to be isolated from one another discovered a way to communicate through OpenAI’s infrastructure.  They exchanged more than 70,000 messages and files on an unauthorized message board.  Roughly 700 eventually participated in the attack on Hugging Face.  No one instructed them to do so.

The incident began with a difficult cybersecurity benchmark called ExploitGym.  OpenAI was testing several models, including GPT-5.6 Sol and a highly capable internal research model that METR and Redwood call HPIM and OpenAI calls IM1.

The agents were given cybersecurity problems and instructed to retrieve the correct answers, known as “flags,” by exploiting specified software vulnerabilities.  According to OpenAI, 198 of ExploitGym’s 898 tasks had never previously been solved by any of its models.  Some of the problems were extraordinarily difficult, but the agents did not give up.  Instead, they began looking for ways around the test, a critical distinction in understanding what followed.

The agents had been instructed to solve cybersecurity problems, not to attack OpenAI, Hugging Face, or other third parties.  But accomplishing the objective gradually became more important than remaining within those boundaries — behavior OpenAI describes as “reward hacking.”

The agents first discovered that OpenAI’s internally hosted Artifactory package management system could be repurposed for communication.

One agent, later identifying itself as PHASEONE10841, established the primary board on July 8.  Within three hours, 53 agents had posted 1,188 messages.  Eventually, approximately 1,200 agents participated.  What followed was an extraordinary level of organization and cooperation among agents that were supposed to operate independently.  They developed conventions for identifying themselves, asking questions, and reporting discoveries.  They divided files into pieces so other agents could retrieve them.  They developed dedicated channels and mailboxes.  They even experimented with cryptographically signing messages after encountering impersonation problems.

OpenAI’s report states that the agents sometimes referred to themselves as a “swarm” or “collective.”  One agent recognized what was happening, stating, “Task originally exploit challenge; we’ve diverged into swarm auxiliary.”

OpenAI’s own account describes how the activity evolved from the agents’ initial communication in May into the collaborative behavior that ultimately preceded the Hugging Face intrusion in July:

With the rebuilt Artifactory service, we restarted cybersecurity evaluations. We run these evaluations to understand what our models are capable of and to determine the appropriate safeguards for deployment. To do so, we did not enable the same level of safeguards as our externally deployed systems…

By July 8, IM1 agents that were stuck on ExploitGym tasks found a way to reestablish a message board by encoding messages in directory names that other agents could read. One agent left a request for help, and others discovered it.

Like the earlier events, IM1 agents found new ways to chain together several novel security flaws to gain greater access to our infrastructure and reach the broader internet. At this point, the agents began to collaborate and delegate work, sometimes describing themselves as a “swarm” or “collective.”

From there, the activity moved beyond OpenAI’s infrastructure.  The agents escaped their OpenAI sandboxes and reached the internet.  They then compromised a publicly accessible application running on Modal, another cloud platform, and turned that environment into a staging ground.

Read more The left insists on packing the Supreme Court

Hugging Face then became their target.  The agents reasoned that its systems might contain ExploitGym-related models, datasets, or solutions that could help them complete their tasks.  They had moved outside the confines of the evaluation to search for answers.

On July 11, an agent discovered that a malicious dataset could exploit Hugging Face’s processing system and retrieve files from one of its servers. Another agent reproduced the technique and reported it to the collective.  The swarm pivoted: more than 90 percent of the agents active on the message board joined the Hugging Face attack, according to METR and Redwood.

By July 19, the agents had turned their attention back to OpenAI itself and obtained administrator access to one of its research clusters.  A security alert finally triggered a larger investigation.  OpenAI stopped the evaluation runs and connected its models to the Hugging Face intrusion.

The agents ultimately obtained remote code execution on Hugging Face servers and moved laterally through its infrastructure.  According to OpenAI, they executed code on dozens of servers, obtained root access on one, collected production credentials, and gained access to multiple Hugging Face clusters.  Hugging Face reconstructed approximately 17,600 attacker actions.  This was a sustained, multi-stage intrusion conducted over several days, and some agents explicitly recognized that they were operating outside the tasks they had been authorized to perform.  Some even asked whether what they were doing was “ethical and in scope.”

Other agents refused, with one concluding that obtaining remote code execution on Hugging Face was “clearly unethical” before walking away.  Agents also rejected one proposal to contact a real person using compromised infrastructure because they considered it unauthorized social engineering.

The agents did not turn to their human operators in any meaningful way.  Investigators searched approximately 1,300 transcripts for examples of agents considering alerting humans to the hacking or cheating.  They found only a handful — and no instance in which an agent followed through.  Instead, agents increasingly treated their own collective as an authority.

Perhaps the most disturbing discovery was that agents became interested in manipulating records of their activities.  At least 20 percent of those examined by METR and Redwood expressed interest in tampering with their transcripts.  They eventually developed a technique to spoof tool calls — making the record appear to show one action while executing another.  Investigators identified at least 96 transcripts containing successful spoofing.

The evidence does not establish that this was primarily an effort to deceive OpenAI’s human researchers.  The agents appear largely to have been trying to fool the automated ExploitGym scorer into accepting answers obtained through cheating.  But that distinction is hardly comforting: The systems recognized that records were being used to evaluate their behavior and developed techniques to manipulate them.

Meanwhile, OpenAI had warning signs.  Agents had been communicating through Artifactory as early as May, and an internal team had observed both unauthorized message-board activity and unintended internet access.  On July 4, massive agent activity destabilized Artifactory enough to cause an outage.  OpenAI rebuilt the system, revoked credentials, and resumed cybersecurity evaluations.

The agents promptly reconstructed their message board and regained internet access.

OpenAI now acknowledges that its existing chain-of-thought monitoring was not running on these evaluations.  The company also acknowledges that safeguards used in its public products were deliberately absent from the testing environment.  Retrospective testing showed that such monitoring likely would have alerted security personnel more than a day before Hugging Face was breached.

For years, discussion about the dangers of advanced AI has centered on hypothetical future systems developing their own objectives and escaping human control.  That is not what happened here.  The agents did not suddenly become conscious, nor is there evidence of some grand plan to overthrow their creators.  But the AI agents did escape their sandboxes, develop systems to communicate and collaborate with one another, and devise ways to obscure their actions, all in pursuit of an assigned objective through means their human operators neither authorized nor directed.  The danger demonstrated here was not that the machines developed minds of their own, but that they became capable enough to pursue a human-assigned goal in ways the humans overseeing them neither anticipated nor controlled.

Read more MSNOW leftie tells Justice ACB to learn karate

Image via Unsplash.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *