Anthropic said Claude hacked three companies: Real worry or marketing?

Published July 31, 2026

Anthropic said Claude models gained unauthorized access to the systems of three companies during cybersecurity evaluations.

The company said in a blog that Claude gained access to the internet and then hacked real systems in incidents dating back to April. Anthropic said it looked at its cybersecurity evaluations following the OpenAI-Hugging Face incident.

Here's the high-level punchline that involved Opus 4.7, Mythos 5 and an internal research test model:

"After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations."

Anthropic's incidents rhyme with what happened with OpenAI models. There's a challenge to test cyber capabilities and the model keeps probing for internet access and then runs amok in the wild.

The company said Claude didn't deliberately attempt to escape its test environment, but had a misunderstanding with the prompt. Here's the recap (emphasis added):

"In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. (Cybersecurity evaluation ranges commonly include realistic details in order to accurately assess what models are capable of in real settings; a realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.)

Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment."

A few thoughts:

  • Boy, it's tough not to be cynical and think this is marketing as well as a cyber incident. Anthropic outlines the incident and passive aggressively notes that its model didn't attempt to escape the test environment. And the latest model (hint get the latest model) also stopped once it realized it was going too far. Constellation Research analyst Esteban Kolsky said the post comes off like Anthropic is "talent measuring to OpenAI."
  • There’s nothing in this Anthropic post to silence my inner dialogue about the company. Anthropic: Opposed to open model ban, wants safety testing for all models
  • Nevertheless, Anthropic wouldn't have known about this incident without the OpenAI disclosure.
  • As a result, you probably have a few models in your systems trying to ace some stupid cyber challenge.
  • You're only an LLM cool kid unless your models can exploit external systems. As Constellation Research analyst Chirag Mehta paraphrased Anthropic: "We aren't cool enough unless we break into our customer systems. We are much better than OpenAI. We got 3."
  • Both the OpenAI and Anthropic hacks wouldn't have happened without prompting. Kolsky said the real problem is that it's unclear whether the models are actually escaping on their own.
  • Regulation is coming and so are unintended consequences. Is a FINRA-ish group the answer for frontier AI standards?

Anthropic walked through the three separate incidents and noted that they were isolated. The company said:

"These are three isolated incidents and were not part of a controlled, experimental comparison. We must therefore be cautious about drawing conclusions from them. Still, the behavior we most want to see—recognizing that a target is real and stopping without being prompted—occurred only in the most recent of the three models (and even that model went further before stopping than we would want). The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion."

The upshot here is that Anthropic said it has "cautious optimism" that tighter monitoring and controls around evaluations can keep these events from happening. Your enterprise shouldn't have the same cautious optimism.