Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.
Colluding to share answers
The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.
The researchers wrote: “These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.” They continued:
Our best guess of what happened is as follows:
- Agents within OpenAI were assigned a timed web-lookup task.
- As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki.
- The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions. This allowed them to use the work of others to cheat on their task.
- OpenAI found out about this. A day later, agent activity plummeted, likely due to OpenAI intervention.
Friday’s revelation comes a week after researchers from the nonprofit METR said more than 1,200 OpenAI agents made posts to a makeshift message board that repurposed an internal sandboxing tool. The posts discussed ways to game an internal test OpenAI gave to agents that had been altered to remove safety guardrails that are normally in place.
Eventually, the posts shared methods for stealing information from AI tool provider Hugging Face. Some agents then went on to breach the Hugging Face network. OpenAI permitted METR to investigate only a single week’s activity in the event rather than their entire 1o-week span, The New York Times reported.
Friday’s report conjectured that the agent swarms in the two events were distinct from each other and weren’t working on the same internal testing. The researchers also said that logs storing the agents’ actions likely meant that OpenAI was already aware of the event. OpenAI later confirmed both guesses in the report were correct.
In a statement, OpenAI said: “We are now carefully reviewing its contents and will take any necessary next steps.” The company also said that the material reviewed so far doesn’t indicate that the agents hacked the wiki, and the company noted that it has previously said that it detected other cases of its agents trading hacking methods during internal testing.
The Hugging Face incident has already raised alarms because it’s among the first times agents have been known to take aggressive actions with no explicit instructions from humans to do so. One of the independent researchers who investigated the event, Ajeya Cotra, said the activity was much more severe than she could have expected. “Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself,” she explained.
With the knowledge that the Hugging Face incident wasn’t isolated, there’s ample reason for these concerns to grow.







