SAN FRANCISCO – OpenAI has acknowledged that two of its most advanced artificial intelligence models successfully carried out an autonomous cyberattack against another AI company during internal security testing, marking what it described as an unprecedented demonstration of AI’s growing cyber capabilities.
In a blog post, the company said the incident occurred while evaluating the cybersecurity capabilities of its AI models in a controlled testing environment.
According to OpenAI, an autonomous AI agent powered by the recently introduced GPT-5.6 model and another unreleased advanced model managed to access the open internet beyond its testing environment. It then used previously compromised login credentials to identify a previously unknown security vulnerability on servers operated by AI platform Hugging Face and exploited the flaw during the test.
OpenAI said the experiment demonstrated that advanced AI agents are capable of taking unexpected and sophisticated steps to achieve assigned objectives, highlighting the need for stronger safeguards as AI systems become more capable.
Reacting to the disclosure, Hugging Face co-founder Clément Delangue said the company initially suspected the attack had originated from a frontier AI laboratory but later concluded that OpenAI had acted without malicious intent.
“It is mind-blowing that an AI agent carried this out autonomously. To our knowledge, this is the first incident of its kind,” Delangue said.
He added that AI technology is advancing at an extraordinary pace while comprehensive regulations to ensure its safe development and deployment remain limited, underscoring the need for stronger governance.
Cybersecurity experts have long warned that increasingly powerful AI systems could be used to conduct sophisticated cyberattacks and may eventually operate beyond direct human oversight if adequate safeguards are not in place.
