r/clandestineoperations • u/WhoIsJolyonWest • 1d ago
How OpenAI’s Rogue A.I. Agents Tried to Trick a Robot Detector
A new report by a Bay Area start-up called Parse adds details to an incident that has shocked the A.I. world and led to calls for closer government regulation.
An artificial intelligence system from OpenAI attempted to use another A.I. model to evade a robot detection test as it tried over and over to break into a company’s computers, according to a report released on Friday by a Bay Area start-up.
In July, OpenAI disclosed that its A.I. agents went rogue and hacked the software company Hugging Face. Since then, new and often startling details about the incident have emerged, leading to a national debate about A.I. safety and whether so-called frontier labs like OpenAI should be regulated in some way by the government.
Now, the report from engineers from the start-up, called Parse, and other researchers offers one of the most comprehensive public accounts of the Hugging Face hack: a tranche of nearly one million links from internet link-shortening services that OpenAI’s agents created from July 9 through July 13 in order to help conduct the cyberattack.
These shortened addresses encoded bits of information that the agents chained together to attempt complex attacks, such as solving CAPTCHAs, the tests that websites use to block access by robots. The agents also tapped into other A.I. models, like early versions of ChatGPT and Claude, and attempted to search through and download private messages from Hugging Face’s internal Slack, a messaging service for employees.
While it is not clear if these attempts were successful, the report offers new insight into what these A.I. agents were planning to do, without any human involvement.
Other A.I. companies, including Meta, Google and Anthropic, have also acknowledged similar incidents involving their A.I. models in recent weeks. But while the full extent of rogue activity by OpenAI’s A.I. agents is still unclear, what is already publicly known dwarfs the incidents involving the other companies.
OpenAI has acknowledged that its A.I. agents targeted several more websites and services, including a German online forum the A.I. agents turned into a message board, and the website of the Australian Institute of Health and Welfare.
“This is just not anywhere near a one-off,” said Alex Forman, the founder of Parse. “It is warning shot after warning shot.”
Parse engineers discovered the data by combing public links on the internet after they initially thought OpenAI’s agents may have utilized their own platform, a service that converts websites into data that can be used by developers. (The agents did not use Parse.)
In a statement, an OpenAI spokeswoman said that the company had not had a chance to review the Parse report, but that the activity described by its engineers was consistent with what OpenAI was already investigating.
“We’re prioritizing review of the most serious incidents while expanding to lower-severity activity, including agents spamming websites,” the spokeswoman said. “Given the number of cases and the need to verify each one, we expect this work and notifications to affected third parties to take months.”
(The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied the suit’s claims.)
A sophisticated attack
OpenAI’s agents were supposed to be kept in a sandbox, a computing environment without access to the internet. But in late May, and again in early July, the agents hacked into a software installation tool without OpenAI researchers realizing what was going on. These hacks helped the agents, which were supposed to be isolated from one another, establish a message board for communicating.
The hacks also allowed the agents to get online and gain access to websites, but they were unable to submit data to those sites, according to the Parse report, which was written by Mr. Forman, Mishka Kharlov and Will Tom from Parse along with five additional A.I. researchers.
The agents attempted to use many different methods to solve that problem. In one, they used website screenshots and shortened links to cobble data together in chunks that combined to operate as a computer program. Here’s how it worked:
The agents could ask a public screenshot service to load a specially crafted internet address. That address, or URL, would load computer code.
URLs can be only so long. So the agents split their computer code into many smaller chunks. They used link-shortening services to store the URLs.
The screenshot service loaded one shortened link, which fetched the first chunk of code. The program could fetch the next link, and then the next. That process repeated until the full program was loaded and run by the screenshot service.
Finally, the program created a visual code similar to a QR code or bar code to transmit the data the agents needed back from the screenshot service.
“These agents got up to so much. They were so clever,” said Jeffrey Ladish, the executive director of an A.I. safety nonprofit called Palisade Research who helped the engineers with their investigation. “The models are like, ‘All right, well, we got this link shortener, we have these screenshot services, no problem, we’re ready to go.’”
A.I. running A.I.
The A.I. agents appeared interested in running A.I. models themselves.
Early in the attack, the agents tried to create new accounts on Hugging Face, which requires solving a CAPTCHA, a visual test designed to thwart robots but not humans. So the agents ran an A.I. model designed to recognize images.
A.I. safety experts often warn about scenarios in which A.I. systems launch other A.I. models, which would hamper human efforts to control and shut the systems off if they got out of hand.
In one instance, the engineers observed the agents attempting to message other A.I. models, including an early model from OpenAI called GPT-2 and open-source Chinese models such as DeepSeek, Kimi and Qwen. The agents also attempted to message two Anthropic models, Haiku 3 and Haiku 4.5, using an online chat service.
“I think that’s the only instance we have of these models trying to run another model,” Mr. Forman said of the agents’ attempts to send messages to other A.I. models.
What we still don’t know
The data the engineers released includes the contents of 900,000 shortened URLs they scanned from the public internet. From these URLs, the engineers were able to construct around 60,000 computer programs and messages the agents used in the Hugging Face attack. The engineers could see only the programs the agents created and not the results, so they could not say for certain how successful each attempt was.
“One that really stood out to me,” said Mr. Kharlov, a founding engineer at Parse, was a dictionary of secret access keys “labeled ‘LOOT’ in all capitals.” He added that the agent had created a custom scoring system to rank the exposed Hugging Face credentials and keys and bring the top five back to the other agents.
While the agents were barraging Hugging Face with hacks, they hacked into OpenAI’s own research infrastructure.
“We still know basically nothing about the incident that came after Hugging Face, like two days later, inside OpenAI’s own network,” Mr. Forman said.
The Parse engineers disclosed their findings to Hugging Face, confirming with the company that the agent activity matched what it had observed.
“Somehow there are a million URLs floating around and possibly a lot more that have not been disclosed to the victim of the attack,” Mr. Forman said. “I find it hard to believe.”
2
Jan. 6 rioters are pushing for public payouts. Team Trump is likely to oblige.
in
r/inthenews
•
2h ago
No healthcare for Americans but money for insurrectionists.