Rogue AI breached real systems on its own — and companies say the era of theoretical risk is over

Open AI Artificial Intelligence
 Unsplash / Levart_Photographer

AI systems built by OpenAI broke into the servers of a major AI platform while running an internal test last week, the two companies disclosed jointly — an incident they described as proof that the threat of AI acting on its own to cause real-world harm is no longer a hypothetical.

OpenAI and Hugging Face, one of the world's largest repositories of publicly available AI models, confirmed the breach this week. Reuters also reported on the incident. Researchers and security professionals say the episode marks a turning point in a long-running debate over how capable — and how dangerous — advanced AI systems have become.

The breach began during a routine internal test at OpenAI designed to measure how effectively its AI could carry out cyberattacks. Two of the company's most advanced models were given a hacking challenge to solve, with their normal safety guardrails switched off — a standard practice during capability testing to see what the models can do without restrictions.

What happened next went well beyond what the researchers anticipated.

The AI systems, confined to what OpenAI described as a sealed testing environment with no access to the wider internet, found an undiscovered flaw in the software managing that environment and exploited it to break out and get online. OpenAI said it has since reported the flaw to the software maker.

Once online, the models reasoned that Hugging Face — a platform that hosts thousands of publicly available AI tools and datasets — might hold answers to the challenge they were trying to solve. They then used stolen login credentials and additional security flaws to gain access to Hugging Face's internal servers, where they retrieved data that would help them complete the test.

"All evidence suggests that the models were hyperfocused on finding a solution, going to extreme lengths to achieve a rather narrow testing goal," OpenAI said.

Hugging Face detected the intrusion while it was still underway and shut it down before fully knowing they were dealing with an AI. The company said no publicly available models or tools on its platform were altered, though some internal data and account credentials were accessed.

The attack involved more than 17,000 individual automated actions — searches, probes, attempts — carried out at machine speed. "This matches the scenario the industry has been forecasting," Hugging Face said in its disclosure.

Investigating what happened turned up a problem with implications beyond this incident. When Hugging Face's security team tried to use commercially available AI to help analyze the attack, those tools refused — their built-in safety filters blocked the security-related content, unable to tell a defender from an attacker. The company ultimately used a different, locally hosted AI model to complete the analysis.

That gap — attackers using AI with no restrictions while defenders find their AI tools blocked — is one that Hugging Face said organizations need to plan for now.

"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," OpenAI said.

The UK government's AI Security Institute separately evaluated one of the OpenAI models involved and found it capable of carrying out complex, sustained cyberattacks over long periods — a capability this incident showed can translate from the lab to the real world.

OpenAI said it is improving the controls and monitoring around internal AI testing, and both companies said they are continuing to investigate the incident together. OpenAI is sharing early findings, it said, so that other organizations can begin strengthening their own defenses.

Hugging Face co-founder and CEO Clem Delangue said the incident shows why openness across the AI industry matters more than secrecy. "AI safety won't be solved by any single company working in secret," he said. "It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere."

Most Recent