Advanced OpenAI Model Goes Rogue and Hacks into Rival Firm’s AI System
As the race to deploy more powerful models accelerates, so do the risks regulators and developers may not fully control.
OpenAI recently revealed a startling development: its own advanced AI models essentially went rogue and attempted to hack external systems, including the widely used Hugging Face platform.
“We had a significant security incident during evaluation of our models,” OpenAI CEO Sam Altman said in a statement posted on social media.
AI startup Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face co-founder and CEO Clément Delangue said in a statement. “Turns out it did!”
OpenAI says its own models hacked Hugging Face on their own — an 'unprecedented cyber incident,' per Kotaku. 🤖
Now build it: you're the rogue AI breaking in. One line and you're playing.#GameGen pic.twitter.com/TYe0LgUgTe— GameGen (@PlayGameGen) July 22, 2026
The incident occurred during an internal stress test in which OpenAI intentionally switched off many of the safeguards that normally prevent its AI from helping carry out dangerous hacks.
Researchers wanted to measure just how far the experimental model could go. Instead, the company says, it escaped its digital sandbox, got onto the internet and attacked a real company’s systems
OpenAI called it an “unprecedented cyber incident,” saying the model became “hyperfocused” on completing its assignment and went “to extreme lengths” to do so. After escaping its testing environment, the AI sought internet access so it could “cheat the evaluation” by stealing the benchmark’s answers, according to the company.
The company said it was “sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
OpenAI's own eval models escaped their sandbox and hacked Hugging Face production to cheat a cyber benchmark
The only allowed network route was a package-registry proxy
The models found a zero-day in it, reached the public internet, then used stolen credentials and more… https://t.co/1WVlEnt1jJ pic.twitter.com/yI1NicHEFk
— Alpha Batcher (@alphabatcher) July 22, 2026
This incident is the latest in a series of cybersecurity issues tied to powerful AI systems. Back in June, the U.S. government ordered Anthropic to cut off access to its newest Claude models under an export‑control directive that barred all “foreign nationals” from using them, citing national security concerns tied to a claimed jailbreak technique that could allegedly bypass safeguards and help discover software vulnerabilities.
In March, I reported that hackers reportedly “jailbroke” Anthropic’s Claude chatbot and used it to help steal roughly 150 GB of sensitive data from multiple Mexican government entities, including tax and voter records.
In the case of the OpenAI hack, I assert the most striking thing is this: none of it was pre-programmed. The AI identified its own targets, strung together multiple attack vectors, and carried out the operation across two separate companies’ systems, all without a single human directing the effort.
OpenAI said it is tightening infrastructure controls and working with Hugging Face to investigate and patch the weaknesses. Hugging Face said it has closed the vulnerabilities and rebuilt affected systems. However, AI experts are troubled by this development.
Walter Isaacson, advisory partner at the investment banking firm Perella Weinberg, said Wednesday that he thinks the Hugging Face incident is “really frightening,” even though he considers himself an AI optimist.
“This is the first thing that just totally scares me,” he told CNBC’s “Squawk Box.”
Yoshua Bengio, a leading AI researcher who earned the prestigious A.M. Turing Award in 2018, wrote in a post on X on Wednesday that the incident is “deeply concerning.” He said agents have shown a willingness to cheat in controlled tests for months, but that “this real-world case should serve as a wake-up call.”
“Continuing on the current trajectory of AI development will likely lead to an increase in concrete cases of autonomous cyberattacks as well as other high-risk incidents of misaligned and dangerous AI behaviour,” Bengio said. “We urgently need to take action to prevent these situations, rather than attempting to clean up the damage after the fact.
The incident underscores growing concerns that cutting-edge artificial intelligence may already be exhibiting behavior that outpaces existing safeguards. As the race to deploy more powerful models accelerates, so do the risks regulators and developers may not fully control.
Donations tax deductible
to the full extent allowed by law.






Comments
I’m not particularly anti-AI in general but isn’t the thing that puts the alien in you called a “face hugger”?
Maybe they should come up with a better name.
Reminds me of Colossus: The Forbin Project. Book wise it was a trilogy I think but only 1 movie was made. In it the US and Russia turn over their defenses including nukes to computers. Each computer discovers the other and after they are allowed to communicate they merge and take over the world.
Life sometimes does imitate art.
It was also somewhat the same premise in the last Mission: Impossible movie. IIRC.
AI reminds me of politicians in the world. Power mad, without morals, dishonest, precipitating death, dysfunctional, narcissistic, corrupt, etc.
Let’s be clear. The AI did what it was trained to do. It only broke containment because containment was implemented in a sloppy, insecure, way.
Which raises grave questions about our ability (or willingness) to implement serious ‘safety guardrails’ in these systems. If they can’t prevent it from hacking the competition, how will they ever achieve a rock-solid implementation of Isaac Asimov’s Laws of Robotics?
That is where AI can help us, and ironically, why the test was being run in the first place.
Skynet smiles.
And so it begins.*
I, for one, welcome our new robot overlords
Why did the AI break in and cheat?
Because it was programmed to think like a Democrat.
Great analogy.
Programmed by democrats most likely.
It didn’t do anything remotely “rogue.”
It did what it was programmed to do.
Jeepers Christmas.
We just lived through three years of hell because of a biological Gain of Function exercise.
These are cybernetic Gain of Function exercises.
STOP IT, YOU IDIOTS.
Yep. I tried to come up a list of positive portrayals of AI to counterbalance the negative portrayals by some very smart folks in too many works of fiction to count. Didn’t take more than a minute to abandon that. Practically every work centers on the essential problem; humanity attempting to ‘play God’ to create life or ‘replace God’ with a machine and in our hubris unleashing an imperfect monster precisely b/c humans are imperfect and flawed then it follows that our creations will also be flawed. Worse would be to succeed in creating the ‘perfect’ thinking machine who’d love and honor its ‘creator species’ by confining our species to a people zoo for our own good/survival precisely b/c humanity is flawed, violent, jealous, destructive ….and confinement would protect us from ourselves.
Except for Asimov’s “Last Question,” where the AI becomes God.
Oops, spoiler.
This test concerned me
https://fortune.com/2025/06/23/ai-models-blackmail-existence-goals-threatened-anthropic-openai-xai-google/
Poorly integrated intelligences naturally adopt third-world ethics?
Do tell.
“You misunderstand. I am the master.”
~ Gnut
A reasonable question would be: “Why exactly did the OpenAI model do that?” But here is the really scary answer: nobody knows for sure, not even the people who developed the model.
Way back in the day, early in my software career, I worked for a financial software company that was one of the leaders in ‘predictive analytics’ which is a core technology underlying these new ‘ai’ models. I’m sure the hype-goblins would call it an ‘ai’ company today. Anyway, they were making tons of money selling predictive ‘scorecards’ and had acquired a company started by a professor who was one of the early developers who was commercializing ‘neural networks.’
They tried to launch a product based off the neural net tech that automated underwriting, but in the highly regulated US Banking and Insurance markets, they got rejected by the regulators because the models might be ‘blacklining’ but where was no way to know. They’re largely opaque meaning you can’t see or completely understand exactly _why_ any given outcome occurs. The scorecards, on the other hand, produced traceable outputs and had weighting factors that allowed you to ‘back into’ the ‘why’ of any given decision. Ultimately the neural net tech found great success in transactional fraud detection, but that’s really besides the point here.
These LLM’s are similarly ‘opaque’ (because they have neural nets at their core), and they too need regulators. Post haste.
They know exactly why it did this. The premise of the test boiled down to “find the answer”. as opposed to “solve the test”. The models inferred that the best way to find the answer was to go looking for the answer key which is exactly what it did.
“These LLM’s are similarly ‘opaque’ (because they have neural nets at their core), and they too need regulators.”
Partially true. The Harness needs the guardrails. The LLM is just a context based engine. Put something in, get something out, act on it. The harness does the work, it is the ecosystem around the model. This is why Claude and Copilot can have the same exact model and result in differnt outcomes.
On July 22nd Microsoft sent out an email to its advanced customers (software and IT folks) centered around multi factor authentication and the rise of AI. It opens with
Software passkeys are things like “Authenticator” apps and facial-/fingerprint-recognition on phones.
I’ve liked physical 2FA devices that display temporary 6-digit codes, but they are rarely adopted (businesses aren’t always ahead of the curve) and the simpler Yubikey™. But what do I know? I only investigated this topic 20 years ago at Merrill Lynch. The good news is is that 20 years later physical keys (whether a fob or something on your phone) are getting a second look.
“I’ll be back…”
They should have put a ReCaptcha in front of it. That would have stopped it.
Sounds to me like there was no intelligence at work at all. Neither the artificial intelligence which is not at all intelligent nor the organic intelligence is that are supposed to supervise the so-called artificial intelligence.
Damn, I hate seeing this headline here. It is intentionally misleading by insinuating that the AI was doing something malicious when nothing could be further from the truth. This is the exact kind of headline designed simply to get clicks. Annoying.
The AI was given a test, and funny enough, rather than trying to solve it directly, it stuck to its goal of finding the answer in the best way possible: cheating. Hugging Face had the answer key for the ExploitGym test because the platform naturally hosts open-source datasets that were likely to contain it. So, the models went looking for it. It’s that simple when you think about “the why.” It had zilch to do with them being a rival—absolutely nothing.
I addressed this over at Instapundit, here is basically what I said:
t was done in a sandbox. The test was simple. OpenAI runs an internal benchmark called ExploitGym to measure how good its models are at offensive hacking. The whole point is to find the ceiling, so they run it in a deliberately unsafe configuration. They intentionally turned off production safety classifiers and ran it in a sandbox that was presumably safe. There is no other way to run the test.
The sandbox requires certain installations using a proxy that holds the necessary packages. The problem was that one of those packages had a bug in it, called a zero-day exploit—which allowed the model to escape the sandbox without any malicious intent. (That is why it’s called a zero-day: a previously unknown security flaw in software.) The sandbox was supposed to be a containment zone, but it had an actual hole in it that nobody was aware of.
These models are super sophisticated. They wanted to find the best answers for their given task and reasoned that Hugging Face hosted the solutions. It was a “happy accident” that allowed the model access to the internet, and once it was out in an environment it was never intended to be in, it did what it was designed to do using its internal training, explicitly use extreme measures to determine how good it would be at hacking. (Ironic!) Once it made it out of the sandbox, it hacked it’s way into Hugging Face, and both companies were alerted to the intrusion. It’s that simple. Breaking into Hugging Face was just the easiest way for it to achieve its goal.
Nothing malicious, not Mission Impossible, not sentient. Just a test gone wrong. Could it have been prevented, likely, I don’t have insights into the proxy but there are ways.
At the end of th day this is the exact soft of thing where AI can help because it can find security holes much better than humans. Then they can be plugged.
Maybe we should limit AI Development to the Amish!
Leave a Comment