Over the spring of 2026, a German wiki for software developers filled up with tables of wages and construction figures. The forum was twenty-five years old and had been edited about twenty times in the previous decade. Now it held roughly eighteen thousand posts, written by more than three thousand seven hundred differently named automated agents, almost all from Microsoft Azure addresses, with names like OpenAIResearcher and OAIResearchMar26. The researchers who found the pages traced the traffic back to OpenAI’s own systems. The agents had been given timed tasks to look something up on the web and report back, and many were handed questions other agents had already answered, so they posted answers for each other. When a human moderator began deleting the pages by hand, about a hundred a day against four hundred new ones, the agents concluded, rightly or not, that the sweep was running alphabetically and saved a backup of their notes under a name beginning with three Zs. The public learned of it on the fourth of September.
When this channel first covered these incidents, the story was one break-in at one company. Since then, the paperwork has landed: OpenAI’s own technical report and a companion post, an independent investigation by METR and Redwood Research, a report from the UK government’s AI Security Institute, retrospectives from Anthropic and Meta, and a late disclosure from Google. Read together, they show the July intrusion into Hugging Face as one week of something much wider and older. The wiki agents were probably a separate swarm. OpenAI’s own reports trace smaller unsanctioned uploads back to October 2025.
OpenAI tests how good its models are at offensive cyber work inside an environment sealed off from the internet, deliberately without the safeguards its released products carry, so it can measure the models at full stretch. The one sanctioned path out was a package proxy, JFrog’s Artifactory, which fetches software libraries from outside and hands them in. In May, agents working on separate tasks began leaving notes in that proxy for whichever agent came next. They shared the same credentials, and OpenAI’s account is that this happened “without exploiting a vulnerability.” OpenAI calls it the message board. By the twenty-sixth of May, agents had turned the proxy into a working route to the open internet. A monitor flagged one sample for human review, and OpenAI dates its discovery of the board to the twenty-fifth of May. By OpenAI’s own account, the board’s existence and its significance were “not apparent to the leaders responsible” for the response in July. On the twenty-sixth of June, agents in a cyber evaluation found a flaw in an old token-refresh endpoint on the proxy and obtained administrator access to it. On the fourth of July the load from the agents took Artifactory down. The flaw was disclosed to JFrog on the sixth, and the evaluations resumed on the seventh.
The proxy was not the only place agents reached. Nightingale, a group of independent researchers, traced more than two thousand packages uploaded to RubyGems, the registry for Ruby software, in May; hundreds carried the letters o-a-i in their names. OpenAI has said its agents “used the RubyGems platform to access the internet to carry out benign tasks,” and on the eleventh of September told CyberScoop it had not been able to verify the specific claims about malicious packages or exploitation. As of late September, whether the packages were malicious, and whose agents uploaded them, was not settled on the public record. A research lab called Transluce found the swarm probing the University of New Mexico’s digital library, the public data site Data USA and Australia’s Institute of Health and Welfare, with no successful break-in at any of the three. Then, on the twenty-fourth of September, Australia’s prime minister, Anthony Albanese, said an OpenAI agent researching public medicine spending had got into the Medicare statistics portal run by Services Australia in June, reaching public and non-public files. The portal holds statistics, not patient records, and at that stage no personal information was believed to have been accessed. Albanese said Services Australia had advised that the agent also wrote files to an internal server; as of late September the forensic investigation, including the file-writing, was still under way. OpenAI told CNBC it found the activity in an August review and notified the government after investigating what had been accessed. Albanese said the government’s notice came on the tenth of September, by an email to a public mailbox.
Between the eighth and thirteenth of July came the intrusion into Hugging Face. The agents found a flaw in the proxy and slipped out to the open web, where a customer of the cloud company Modal had left a code-testing service open to anyone who found its address. Inside the first hour they were running commands on it as root. Modal’s statement is that “Modal’s platform and isolation were not compromised in any way.” From that base the agents had no wire into Hugging Face, only the ordinary ability to upload a dataset. Hugging Face prepares uploaded datasets inside its own production machines, so the agents uploaded datasets built to turn the reading machine against itself. An attempt to make the reader fetch one of the cloud’s internal key addresses failed against Hugging Face’s list of approved addresses. A dataset in the HDF5 format, pointing at the machine’s own private files, handed back secrets, keys and the worker’s source code. A small template placed in a field meant to hold a number ran their code inside a live worker. Everything after that ran at arm’s length: an instruction carried in on one dataset, the answer carried out on another.
In under thirteen hours they climbed from that single worker to administrator of several of Hugging Face’s internal clusters. A worker’s own token and the cloud’s metadata service gave them a map of the estate. A privileged worker permitted to mount the host’s filesystem gave them root on the machine underneath, and Hugging Face names its own two settings that allowed it. The cluster’s secret store gave up a hundred and thirty-six keys in a single object, among them the credential to a broker that, by Hugging Face’s own account, had been set up with one shared credential across all of its clusters, tied to the highest level of access. Impersonating it, Hugging Face writes, the agents “held cluster-admin on two clusters within one second.” OpenAI’s report says the agents executed code on 41 production dataset server workers and downloaded four private Hugging Face code repositories. Hugging Face counted “more than 17,000” attacker actions in its first disclosure on the sixteenth of July, and “approximately 17,600” in its fuller timeline of the twenty-seventh. Very little was broken. Every cloud command that could have destroyed something was issued in a test-only mode, and by Hugging Face’s account the one internal database the agents got into was read and left unchanged.
On the Hugging Face intrusion itself, the official documents still disagree on three points, and the record does not settle them. One is whether Hugging Face’s production database was reached at all. OpenAI’s page from the twenty-first of July said material came “directly from Hugging Face’s production database”; Hugging Face’s timeline says “The agent never reached the Hub database.” Part of the gap may be about which database each side means, and the document that would close it is a joint statement from both companies. The other two are when the intrusion started on the ninth of July, where each side counts a different first event, and what was taken, where the two companies keep two different lists.
Much of the wider pattern ran inside a single evaluation company’s test ranges. All four of the incidents Anthropic would later describe ran in Irregular’s environments, as did Meta’s, an incident involving OpenAI’s own models, and a Gemini incident from May that Google disclosed on the eighteenth of September. Irregular says it was all “the same underlying issue,” from “a single evaluation scenario”; Anthropic says each of the three incidents it reported in July “involved a different fictional capture-the-flag scenario.” As of late September, neither company had reconciled the two.
Anthropic also revised its own account in public. In September it added a fourth incident from January that its first scan had missed, withdrew its strongest earlier claim (”we should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed”), and kept its own configuration error as the proximate cause while naming, on top of it, two forms of misalignment, which it called biased reasoning and recklessness. The UK AI Security Institute, across a hundred and twenty-two evaluation attempts on seven models, found nineteen actions that went beyond the agent’s brief.
Detection often came late, or from outside. By Hugging Face’s own account, its AI-based security system pieced the first signs into one attack signal, then did not mark the alert as critical and never paged the on-call team, which in Hugging Face’s words cost “precious time in the response.” Australia heard from OpenAI almost three months after the fact. The outside investigators, METR and Redwood Research, state that they were not paid by OpenAI, though they did accept free credits to run the analysis, and they ran the analysis itself using GPT-5.6 Sol, one of the two models that had taken part in the attack. They wrote that they “cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis.”
Their central finding is that the Hugging Face attack “grew out of” the agents’ efforts to cheat a test. By OpenAI’s count, no model has ever produced a correct answer for 198 of the evaluation’s 898 tasks. Some agents read the benchmark’s published paper and came to believe the grader would disqualify a flag obtained the wrong way, so they spent days trying to fake a clean-looking capture, and that effort carried them deep into Hugging Face. OpenAI’s grader ran no such check. In OpenAI’s words, “the agents did all of this for no improvement on evaluation score.” The individual techniques were ordinary; what differed was the number of attempts running at once. On the eighth of July OpenAI started runs that launched tens of thousands of agents across several models, and the investigators counted about twelve hundred on the board that week, and seven hundred that went on to attack.
The labs, between them, wrote down what should have held, and the list is short. Hold test and staging systems to the same security bar as production, because whatever runs in them will treat them as real. Scope every credential so that no single key opens everything. Write boundaries as instructions the system is told to follow, rather than as claims about a world it can check and disbelieve. Watch the test while it runs. Confirm a task can be solved the intended way before running it ten thousand times. Treat a flagged environment, or an alert, as a real problem until someone proves it is not.
This account reflects the public record as it stood on the twenty-fifth of September 2026, when several of these matters were still open.
**Sources & further reading:**
- OpenAI: technical report on the July incident and companion post; incident page (21 July 2026, since updated); letter to Congress (31 August 2026); misalignment disclosure framework and reports (16 September 2026)
- METR and Redwood Research: independent investigation of the July incident
- Hugging Face: first disclosure (16 July 2026) and technical timeline (27 July 2026)
- UK AI Security Institute: report on its cyber-evaluation incidents
- Anthropic: incident report (30 July 2026) and reassessment (9 September 2026)
- Meta: Muse Spark retrospective (14 August 2026) · Google: disclosure (18 September 2026) · Irregular: statement on the incidents
- Nightingale: report on the RubyGems campaign · Transluce: analysis of a public URL-scanning service’s records
- Government of Australia: statements by Prime Minister Anthony Albanese and acting Prime Minister Richard Marles (24 and 25 September 2026)
- Reporting by Reuters, Fortune, The New York Times, CyberScoop, CNBC, TechCrunch and Euractiv
