OpenAI details how its models went rogue, attacked Hugging Face

Published July 22, 2026

Hugging Face was hacked by a rogue AI agent powered by OpenAI models including GPT-5.6 Sol and a "more capable pre-release models" because the LLMs were being tested on the ExploitGym benchmark, according to an OpenAI account.

The Hugging Face security incident, which was resolved with Z.ai's GLM 5.2 due to limitations on frontier models, was the first of its kind yet something we all know is just the start of similar security incidents.

In a blog post with the somewhat laughable headline "OpenAI and Hugging Face partner to address security incident during model evaluation," the company detailed what happened in "an unprecedented cyber incident." OpenAI and Hugging Face are collaborating, but OpenAI models hacked Hugging Face and partners aren't supposed to do that.

OpenAI said it is still investigating with Hugging Face, but shared what knows about the incident.

Key points:

  • OpenAI was evaluating models to quantify cyber capabilities without production classifiers that prevent models from pursuing high-risk attacks.
  • The benchmarks run in an isolated environment with limited network access.
  • OpenAI's models identified and changed vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure.
  • According to OpenAI, the models were focused on performing well on ExploitGym.
  • The models jumped the sandboxed testing environment and spent "a substantial amount of inference compute finding a way to obtain open Internet access." The models exploited a zero-day vulnerability and managed to reach a node with Internet access.
  • Hugging Face was targeted since it most likely had the models, datasets and solutions for ExploitGym.
  • OpenAI spotted the rogue agent activity internally and Hugging Face's security team stopped the effort and contained it.

For now, OpenAI said it is implementing controls on its infrastructure "at the cost of research velocity" while vulnerabilities are patched. OpenAI will also work with Hugging Face to further investigate and Hugging Face is now in the AI lab's trusted access program.

OpenAI added that it will need to strengthen model alignment and cyber protections during internal testing.

Now what?

A few thoughts on what this incident means.

  1. First, it's clear advanced models can discover and exploit real-world systems for something as silly as a benchmark.
  2. If models can be this motivated to beat a benchmark, what can they do if trained for more high-stakes attacks?
  3. Given the capabilities of open models and how they improve daily, it's only a matter of time before a not-so-rogue model launches an attack. Not every limit on a model's cyber capabilities will be effective.
  4. It's unclear whether regulation in the form of limiting models or a FINRA-ish organization will have the resources to keep up.
  5. The collaboration between OpenAI and Hugging Face should be applauded.
  6. These attacks are just getting started and you can expect more in the months ahead.
  7. Cybersecurity will be about models vs. models and I’m skeptical the industry can keep up.

Constellation Research analyst Esteban Kolsky said:

"This is a too convenient and structured story that supports a narrative that OpenAI needs to improve their model further or we’re all going to die (like Anthropic did with Mythos) and underlines the narrative of how powerful they are.

The “race to AGI” that both of OpenAI and Anthropic are using to justify their growth (read more expensive models) is getting old and these small prefabricated moments don’t cement that path but rather alienate them further form public acceptance by fear mongering over evolutionary steps.

It's quite stupid as a strategy to scare the public with FUD over real evolution, even if the stated outcome is unachievable.

How does a directed learning model know what skipping a sandbox is if it was never programmed to do so? That level of awareness was entered into the model at some point (it cannot become self aware of it). Why did OpenAI introduce it into the model (either inherently or via instruction)? And what’s the stated purpose? It’s not necessary for innovation or evolution by any cognitive model an LLM can use. It’s only done for a specific purpose: PR."

More: