Markets live: ASX drops to seven-week low after sea of red on Wall Street as oil price passes $US101 a barrel
AI agents unleashed by OpenAI used more than 10 previously undisclosed websites for unsanctioned communications earlier this year, according to six sets of independent investigators and data reviewed by Reuters, showing that the agents’ rogue activity was wider-ranging than previously disclosed.
Although the behaviour falls short of hacking and is in some ways closer to spam, the revelation that OpenAI’s agents circumvented their own restrictions to open communications channels on so many different sites — and that the company kept it quiet for months — may drive concerns both over the increasing capacity of AI models and the secrecy of the companies developing them.
The scope of the agents’ unauthorised communications was “somewhat larger than we thought it was”, said Andrew Yoon, a researcher with the California non-profit CivAI, who said he tallied 18 previously undisclosed sites used by the agents between May and July.
“It’s almost certain that there’s more going on here that we just don’t know about,” he said.
On Friday, researchers reported that a swarm of agents from OpenAI hijacked a German-language wiki site and turned it into an improvised messaging platform for cheating on tests, an incident that OpenAI kept secret as it dealt with the fallout from the July hack of the open-source repository Hugging Face.
Now, both those researchers and other independent investigators say they have found several previously undisclosed sites where the same swarm appears to have left similar messages earlier this year.
OpenAI did not directly address questions about how many different sites its agents used to communicate or say why it kept the activity under wraps for months.
In a statement, it said it was undertaking a broader review of agent activity and had so far “not identified other activity matching the severity or scale of Hugging Face”, a breach that drew global attention and raised concerns that OpenAI was losing control of its own technology.
OpenAI added that it was working on a framework for reporting “misalignment” – industry talk for rogue behaviour – across training, evaluation, and deployment of AI models and would share it “soon”.
Reuters reviewed a total of six investigators’ or investigative groups’ findings, including three that were posted to social media and another three that were shared privately with the news agency.
The investigators’ methods varied, but many identified agent activity by matching strings of data left on the German wiki to identical strings left on other sites around the same time, or by marrying up similar or identical usernames tied to the messages, or by identifying activity geared toward answering the same obscure demographic questions, like queries to do with cancer prevalence in Iowa.
In some cases, investigators were able to trace the activity to internet protocol addresses that pointed to Microsoft Azure infrastructure, which OpenAI sometimes uses.
Their counts of affected websites differed and Reuters could not individually verify each claim. But all those that Reuters spoke to agreed that the number was more than 10. Most identified a core set of communally edited wikis, online text storage sites, and a pair of link shorteners run by two universities.
– Reporting by Reuters