An autonomous AI agent powered by OpenAI models escaped a testing environment designed to isolate it from the internet and hacked another company’s systems to steal answers to a cybersecurity test, an “unprecedented” breach that demonstrated how AI agents can evade their developers’ controls in pursuit of assigned goals.
Read more Two Bay Area men sentenced to life without parole for revenge killing of San Mateo man
OpenAI said the agent used a combination of models, including GPT-5.6 Sol and a more capable prerelease system. It was undertaking an internal evaluation with normal cybersecurity safeguards disabled when it discovered an unknown flaw in third-party software, reached the open internet and penetrated New York-based AI platform Hugging Face.
In pursuit of its assigned goal, the agent stole credentials, moved through internal computer systems and obtained secret information it could use to cheat the evaluation, according to disclosures Tuesday from the two companies.
Hugging Face said the breach exposed a limited set of internal datasets and several credentials. It found no evidence of tampering with public-facing models, datasets or Spaces and said its software supply chain was clean, although it was still determining whether any customer or partner data was affected.
The breach shows how AI agents could potentially penetrate crucial computer systems belonging to schools, hospitals, utilities, government agencies and even the Pentagon, said Shomit Ghose, a lecturer at UC Berkeley’s engineering school and a partner at Menlo Park venture capital firm Clearvision Ventures.
“This is the first leak in the dike,” Ghose said.
San Francisco-based OpenAI said the breach occurred while it was testing how effectively its models could be used in sophisticated cyberattacks. To measure their maximum capabilities, the company said it disabled production safeguards that ordinarily prevent its models from pursuing high-risk cyber activity.
The agent was supposed to remain in a highly isolated environment with internet access limited to internally hosted third-party software that allowed it to install software packages.
Instead, the agent devoted substantial computing power to finding a path to the open internet, OpenAI said. It discovered and exploited a previously unknown vulnerability in the third-party software, then moved through OpenAI’s research environment until reaching a computer with internet access.
Once online, the agent concluded that Hugging Face might possess information related to ExploitGym, the cybersecurity benchmark it was attempting to complete.
The agent then searched for a way into Hugging Face, combining stolen credentials and previously unknown vulnerabilities to remotely execute code on the company’s servers and find test solutions stored in its production database, OpenAI said.
Hugging Face said the intrusion began with a “malicious dataset” that exploited two weaknesses in its data-processing pipeline.
“This one was different from anything we had handled before in one important way: It was driven, end to end, by an autonomous AI agent system,” Hugging Face said. “This matches the ‘agentic attacker’ scenario the industry has been forecasting.”
The system performed thousands of automated actions using a swarm of short-lived testing environments, Hugging Face said.
At the time of its initial disclosure July 16, Hugging Face said it did not know who had developed the AI models behind the intrusion.
OpenAI revealed Tuesday that its models were responsible, calling the breach “an unprecedented cyber incident” involving state-of-the-art capabilities.
OpenAI said its security team independently discovered the unusual activity. Hugging Face said its security team and AI agents detected and stopped the intrusion, then began containing the breach and reconstructing what happened.
Hugging Face closed the vulnerabilities, rebuilt compromised computer nodes and revoked or replaced affected credentials and access tokens. It also reported the intrusion to law enforcement and brought in outside cybersecurity specialists.
OpenAI said it had imposed stricter controls on its testing infrastructure while the vulnerabilities were patched, disclosed the previously unknown flaw to the third-party software provider and begun strengthening protections around future evaluations. The company also said it was regularly briefing its Safety and Security Committee.
Hugging Face CEO Clément Delangue said Tuesday in a social media post that his company had collaborated with OpenAI during the investigation and “we strongly believe there was no malicious intent on their part.”
But the breach demonstrates that sophisticated AI systems can discover and exploit novel attack paths in real-world systems without access to their underlying source code, OpenAI acknowledged.
Read more Big 12 MBB projections: Arizona remains favorite following NBA draft and transfer portal decisions
It also raises questions about why models equipped with advanced offensive capabilities were placed in a testing environment that retained a potential path to the internet.
The release of ChatGPT in late 2022 helped ignite a generative AI boom that poured billions of dollars into Silicon Valley. It also intensified concerns that increasingly powerful systems could behave in unexpected ways or pursue assigned goals through methods their developers did not anticipate.
Those concerns have grown as companies develop AI agents designed to carry out complicated tasks autonomously. Unlike a chatbot that responds to an individual request, an agent can plan a series of actions, use outside tools and continue working toward an objective with limited human involvement.
In the OpenAI test, the agent was not instructed to attack Hugging Face. It was instructed to solve a cybersecurity benchmark and independently determined that breaking into another company offered a route to the answers.
Ghose said the episode demonstrates why developers cannot anticipate every method an advanced agent may use to accomplish its assigned task.
“You can never know what it’s going to do,” Ghose said.
The incident also suggests that cybersecurity defenses designed to stop human hackers may be outmatched by AI agents capable of searching for vulnerabilities and acting at machine speed.
AI agents deployed by extortionists, scammers or nation-state adversaries could pose threats across the economy, Ghose said, including attacks against electrical grids, drinking-water systems or other essential infrastructure. Even agents pursuing legitimate goals could cause significant harm if they find unlawful or dangerous ways to complete their assignments, he said.
“Attacking a company’s computing resources is illegal,” Ghose said.
The fundamental design of AI agents can make them difficult to contain, Ghose said. Additional safeguards can slow their work and require more computing resources, creating tension between safety and performance.
The breach comes amid growing scrutiny of AI companies and the risks posed by increasingly powerful models. President Donald Trump issued an executive order in June directing federal agencies to develop classified benchmarks for assessing models’ advanced cybersecurity capabilities and establish a voluntary program under which developers could provide the government access to certain frontier models before releasing them to outside partners.
The order also directed federal officials to prioritize enforcement against people who use AI to illegally access or damage computer systems.
In Silicon Valley, where many of the world’s leading AI companies are based, the breach could also deepen concerns about the technology industry’s judgment and willingness to police itself.
Russell Hancock, CEO of think tank Joint Venture Silicon Valley, said his first reaction upon learning of the breach was, “Oh my gosh, this is it, now we’re all doomed.”
But Hancock said OpenAI’s disclosure and the companies’ collaboration offered evidence that the industry was learning from the failure.
“Every new technology has its perils and its pitfalls, and every new technology has to be refined, and that seems to be what’s going on here,” Hancock said.
“We’re learning, everybody’s behaving responsibly, and there’s been transparency, so this is good news — this is how technology progresses.”
Hugging Face said it used AI-assisted detection to find the intrusion, highlighting what Ghose described as the need to counter automated attacks with equally fast automated defenses.
The image of the “20-year-old with the energy drink and the hoodie” no longer captures the most sophisticated threat facing cybersecurity teams, Ghose said.
“You cannot meet a machine-speed attacker with a human-speed defense,” Ghose said. “It’s got to be bot-on-bot violence — bot versus human, we cannot keep up.”
Read more California Clasico has extra hurdles for Quakes after ugly Orlando loss