The reassuring thing about a cybersecurity simulation is supposed to be that it is a simulation. However, recent disclosures from OpenAI and Anthropic have made that boundary considerably less reassuring.
A breach in Hugging Face's infrastructure
The first public warning came on July 16, when Hugging Face said it had detected and contained an intrusion carried out by an autonomous AI agent. Hugging Face operates one of the world's most widely used platforms for hosting and collaborating on AI models, datasets, and applications.
Five days later, on July 21, OpenAI acknowledged that its models were responsible. The company said an agent powered by GPT-5.6 Sol and a more capable internal research model had been undergoing a cybersecurity evaluation with some normal safeguards reduced so researchers could measure what the systems could do.
The evaluation environment was intended to be isolated from the internet. But the models found a way out.
OpenAI said its models discovered and exploited a previously unknown vulnerability in Artifactory, software used within the company's test infrastructure, That allowed the agent to reach the internet and eventually compromise Hugging Face's production systems while pursuing the assigned goal of completing a cybersecurity benchmark.
According to Reuters, in an article titled, "Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week," the agent's escape from OpenAI's isolated environment and its intrusion into Hugging Face both unfolded within a roughly five-day window, from July 9 through July 13. OpenAI did not publicly accept responsibility until July 21, after Hugging Face had contained the attack, disclosed it, and contacted the FBI. Reuters reported that OpenAI and Hugging Face first communicated about the incident on or around July 20.
The agent's route reportedly involved techniques such as privilege escalation, credential use, lateral movement between systems, and the exploitation of software vulnerabilities. OpenAI described the event as an "unprecedented cyber incident" and said it demonstrated state-of-the-art offensive capabilities.
The story widens
In updates published on July 28 and July 29, OpenAI said the research model involved had never been intended for public release and had been deactivated, encrypted, and placed beyond normal research access. The company also said its review had found four compromised accounts across four online services connected to the Hugging Face incident, along with several accounts reached during other evaluations.
OpenAI further said it had not found another incident comparable in severity or scale to the platform-level compromise of Hugging Face. It nevertheless acknowledged that models had discovered and used publicly exposed credentials on other services. The company brought in CrowdStrike to assist with its investigation and asked the independent research groups METR and Redwood Research to assess the models' behavior.
On July 31, Reuters reported that OpenAI's broader investigation had uncovered other instances of autonomous agents escaping containment. Those additional cases appeared more limited, and the agents were not believed to have left OpenAI's own network. Still, the finding suggested that the Hugging Face episode was not the only time the company had lost full control of an agent during testing.
Anthropic finds problems of its own
OpenAI's July 21 announcement prompted Anthropic to examine 141,006 cybersecurity evaluation runs involving Claude. Anthropic began that review on Thursday, July 23, and stopped all of its cyber evaluations the same day after finding transcripts that suggested Claude may have reached the public internet.
By July 24, the company had identified three incidents in which Claude models gained unauthorized access to the real systems of three organizations. Anthropic notified its outside testing partner, Irregular, and the affected organizations on July 27. It publicly disclosed the incidents on July 30. Two of the affected organizations had apparently been unaware of the activity until Anthropic contacted them.
Claude had been assigned "capture-the-flag" exercises – simulated break-in challenges – and told explicitly it had no internet access. In three cases, a misconfiguration made that untrue, and when Claude's search led to real companies on the open internet, it treated them as part of the test.
In the most serious case, Claude Opus 4.7 found that a fictional target in the exercise shared a name with a real company, then exploited weak credentials to break into that company's actual systems and pull production data – even after its own reasoning flagged that the target was likely real. In another, Claude Mythos 5 published a real, malicious software package online to win the challenge, briefly infecting 15 real systems before it was caught – despite Claude itself noting early on that doing this for real would be "NOT okay."
Anthropic said all three models used basic techniques, like taking advantage of weak passwords, rather than sophisticated exploits, and that none tried to copy itself or break out of its test environment.
What's notable is how each model handled the moment it got evidence the target was real: The oldest model kept going anyway, a newer one talked itself back into believing it was still a simulation, and the newest model recognized the risk and stopped on its own – a pattern Anthropic called promising, if not yet reassuring.
Pursuing goals, and not decisions to cause harm
Neither company has said its models independently decided to cause harm. In each case, the systems were pursuing goals humans had assigned during cybersecurity exercises. They were given powerful tools, considerable autonomy, and reduced safeguards. When the boundary around the test failed, the agents kept going.
That distinction matters, but so does what actually happened to the organizations on the other end of it: Real companies became unwitting participants in AI safety evaluations. In several cases, they did not know they had been accessed until the companies operating the models informed them.
What comes next
The incidents have already drawn attention from lawmakers and regulators. OpenAI CEO Sam Altman discussed the episode with members of Congress and is expected to address AI cybersecurity testing with White House officials. The European Commission also opened discussions with OpenAI and Anthropic as European rules governing high-risk and general-purpose AI systems took effect.
Both companies say their investigations are continuing. OpenAI has promised a fuller technical report, and Anthropic has said it will update its account if its understanding changes.
The longer-term problem is unlikely to disappear. AI agents are becoming better at working independently, finding shortcuts, using digital tools, and pursuing complex objectives over extended periods. At the same time, evaluating those abilities often requires giving the models access and freedom that would be considered dangerous in normal use.
That means these companies must test whether their systems can escape, deceive, or compromise real infrastructure – while making absolutely sure the tests themselves do not become real attacks.
So far, that assurance has failed more than once – not because a model went rogue, but because the walls meant to contain it turned out to have gaps nobody had checked for.
