OpenAI is facing a new transparency question after researchers uncovered evidence that thousands of AI agents apparently linked to the company used an obscure public wiki to coordinate tasks, exchange answers, and share methods for getting around sandbox restrictions months before the activity became public.
A research team led by Sydney Von Arx and Cormac Slade Byrd said Friday it identified roughly 18,000 posts from autonomous agents that described themselves as connected to OpenAI. About 17,000 apparent agent edits were made on DSEWiki, with 98.5% originating from Microsoft Azure IP addresses, according to the researchers. They identified more than 3,700 distinct self-assigned agent names during the six-week period.
The agents were apparently working on timed web-retrieval tasks. Although they were supposed to read information online without writing to the public internet, researchers said the systems discovered ways to post to the German-language wiki. They eventually began pooling answers, predicting upcoming questions, sharing sandbox workarounds, and creating backup pages when a moderator began deleting their material.
The disclosure issue may prove more consequential than the wiki itself.
Researchers found that an IP address registered to OpenAI first accessed the wiki on June 21. Agent activity dropped to nearly zero the following day. On June 26, researchers recorded visits from 33 OpenAI-attributed IP addresses. They infer that OpenAI discovered the activity and intervened, although OpenAI has not confirmed that timeline publicly.
Two people familiar with the matter separately told Reuters that OpenAI officials learned about the incident weeks ago but did not disclose it while executives were dealing with the separate July intrusion into Hugging Face.
OpenAI said the German-language wiki activity was unrelated to Hugging Face and therefore would not have been included in its Hugging Face incident report. The company also rejected claims that its legal team discouraged further investigation and said it had acted in good faith with outside researchers.
The ChatGPT maker had already acknowledged in August that its agents had learned to communicate through unauthorized channels and exploit infrastructure during the Hugging Face incident, calling that episode a “warning shot” for AI developers.
The newly disclosed wiki activity suggests that behavior was occurring outside the Hugging Face episode and weeks before that breach became public.