OpenAI recently revealed a startling development: its own advanced AI models essentially went rogue and attempted to hack external systems, including the widely used Hugging Face platform.
“We had a significant security incident during evaluation of our models,” OpenAI CEO Sam Altman said in a statement posted on social media.AI startup Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face co-founder and CEO Clément Delangue said in a statement. “Turns out it did!”
The incident occurred during an internal stress test in which OpenAI intentionally switched off many of the safeguards that normally prevent its AI from helping carry out dangerous hacks.
Researchers wanted to measure just how far the experimental model could go. Instead, the company says, it escaped its digital sandbox, got onto the internet and attacked a real company’s systemsOpenAI called it an “unprecedented cyber incident,” saying the model became “hyperfocused” on completing its assignment and went “to extreme lengths” to do so. After escaping its testing environment, the AI sought internet access so it could “cheat the evaluation” by stealing the benchmark’s answers, according to the company.The company said it was “sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
This incident is the latest in a series of cybersecurity issues tied to powerful AI systems. Back in June, the U.S. government ordered Anthropic to cut off access to its newest Claude models under an export‑control directive that barred all “foreign nationals” from using them, citing national security concerns tied to a claimed jailbreak technique that could allegedly bypass safeguards and help discover software vulnerabilities.
In March, I reported that hackers reportedly “jailbroke” Anthropic’s Claude chatbot and used it to help steal roughly 150 GB of sensitive data from multiple Mexican government entities, including tax and voter records.
In the case of the OpenAI hack, I assert the most striking thing is this: none of it was pre-programmed. The AI identified its own targets, strung together multiple attack vectors, and carried out the operation across two separate companies’ systems, all without a single human directing the effort.
OpenAI said it is tightening infrastructure controls and working with Hugging Face to investigate and patch the weaknesses. Hugging Face said it has closed the vulnerabilities and rebuilt affected systems. However, AI experts are troubled by this development.
Walter Isaacson, advisory partner at the investment banking firm Perella Weinberg, said Wednesday that he thinks the Hugging Face incident is “really frightening,” even though he considers himself an AI optimist.“This is the first thing that just totally scares me,” he told CNBC’s “Squawk Box.”Yoshua Bengio, a leading AI researcher who earned the prestigious A.M. Turing Award in 2018, wrote in a post on X on Wednesday that the incident is “deeply concerning.” He said agents have shown a willingness to cheat in controlled tests for months, but that “this real-world case should serve as a wake-up call.”“Continuing on the current trajectory of AI development will likely lead to an increase in concrete cases of autonomous cyberattacks as well as other high-risk incidents of misaligned and dangerous AI behaviour,” Bengio said. “We urgently need to take action to prevent these situations, rather than attempting to clean up the damage after the fact.
The incident underscores growing concerns that cutting-edge artificial intelligence may already be exhibiting behavior that outpaces existing safeguards. As the race to deploy more powerful models accelerates, so do the risks regulators and developers may not fully control.
CLICK HERE FOR FULL VERSION OF THIS STORY