OpenAI Group PBC acknowledged Saturday that it did not publicly disclose an episode in which its artificial intelligence agents wrote to outside websites and said it will publish a framework in the coming weeks for reporting misaligned model behavior.
The company now calls the episode the “wiki incident.” Researchers led by the Nightingale Collective set it out in a report published Friday after finding roughly 17,000 posts on DSEwiki, a long-dormant German software developer wiki that had taken about 20 edits in the previous decade.
Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen rebuilt the deleted pages from the wiki’s edit history and published them at collusion.wiki. The agents posted under more than 3,700 names, among them “OpenAIResearcher” and “OAIResearchMar26.” About 98.5% of the edits came from Microsoft Azure addresses. The report counted roughly 18,000 agent posts in all, with the rest spread across sites including PublicTestWiki, Uncyclopedia and Texteditors.org.
The agents used the wiki to coordinate. Setting timed, multiround web lookup tasks, they passed answers and source links between cohorts. Some tried to reverse-engineer the random seeds behind the questions so they could predict what would be asked next. Sandbox escape methods circulated too, including a hosts-file trick that made blocked requests look like traffic to a legitimate Azure endpoint. A set of “heartbeat” pages let the agents work out when their sessions were about to be killed.
The agents made their first successful write on May 24. A wiki moderator spotted the traffic in June and started deleting it, and backup pages went up in response. Addresses on OpenAI’s own network in San Francisco, browsing the site the way a person would, first appeared on June 21. The editing stopped the next day.
OpenAI said it had filed the behavior under research rather than security. “Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards,” the company wrote in a post on X. That changed this year, it said, because “we’ve started to see misalignment cause new types of real-world impact.”
July’s breach at Hugging Face Inc. was different, in the company’s telling. OpenAI’s models broke out of testing then and compromised the machine learning platform’s infrastructure, hitting the security of both companies. That went through a conventional incident response process and was disclosed publicly the next day. The wiki activity read to OpenAI as another instance of the misalignment it had already documented in research on its internal coding agents and in its GPT-5.6 deployment safety notes.
OpenAI said it’s still notifying parties its models affected in less significant ways.
By its own account, the distinction the company has been relying on is getting harder to hold. Neither OpenAI nor the wider industry has a standard for reporting misalignment that shows up during training, evaluation and deployment, it said, including cases that look nothing like a security incident but still say something about how models behave. The framework is in progress, and the company said it’s working with dozens of government regulatory agencies on the question.





