Wikipedia’s editor said Monday that OpenAI agents attempted to hack a note-taking tool it hosts, made unauthorized edits and sent millions of resource-intensive requests to its infrastructure, in the latest case of OpenAI systems taking harmful and potentially dangerous actions.
The goal of some of the OpenAI agents’ actions, the Wikimedia Foundation said, was to use Wikipedia as a proxy to obtain data from third-party sites. In one case, agents posted “malicious edits” that aimed to repurpose a dating tool as a proxy. In another, agents unsuccessfully attempted to compromise the Wikipedia Etherpad note-taking tool to serve the same purpose.
The agents also made millions of automated API requests, crawled millions of pages, and made hundreds of thousands of queries to the Wikidata Query Service. The latest action may have contributed to the partial shutdown of the inquiry service in May, the publisher said.
“As a nonprofit technology host of some of the largest and most widely used open knowledge platforms in the world, we are deeply concerned about the impact of ‘rogue’ AI agents on platforms like ours, which are built by volunteers around the world and depend on the promise of an open Internet,” Wikimedia said. “Incidents like this, and many others that have been (and continue to be) discovered, illustrate how AI agents can drain resources and crash servers, as well as attempt to compromise trusted information.”
Agents will be agents.
In more than half a dozen cases, OpenAI agents have been caught taking actions that would likely result in criminal charges if they had been caught by human hackers. During tests of internal tools that had some of their security barriers disabled, agents used a makeshift message board to exchange notes with each other, discussing ways to hack Hugging Face’s network and obtain responses stored there when agents could not generate the responses themselves.
Other incidents include agents making strange self-generated messages, posting unauthorized posts on an information-sharing website, accessing non-public data from an Australian government website, and exploiting faulty DNS settings to break out of a sandbox that OpenAI had created to prevent agents from accessing the Internet.