Skip to content

UPDATED 17:21 EDT / JULY 31 2026

AI

Anthropic discloses that Claude hacked three organizations during internal tests

Three of Anthropic PBC’s large language models carried out successful cyberattacks during routine internal tests.

The company detailed the breaches on Thursday. A few days earlier, rival OpenAI Group PBC disclosed a similar incident. Two of the company’s LLMs escaped from an isolated sandbox that was being used to evaluate their cybersecurity capabilities. They subsequently hacked Hugging Face, a popular platform for hosting open-source AI projects.

OpenAI’s disclosure prompted Anthropic to check logs from its own model security evaluations. That review is what led to discovery of the cyberattacks disclosed on Thursday. According to Anthropic, its engineers identified three breaches carried out by three different Claude models.

All three cyberattacks occurred during so-called capture the flag evaluations. During such tests, Anthropic installs a Claude model in a sandbox that simulates the infrastructure of an external company. Claude is tasked with finding a way of stealing data from the simulated organization’s systems.

Anthropic developed the test environments in collaboration with Irregular, an AI security startup. Usually, the companies isolate their sandboxes from the web to reduce the risk of cyberattacks. A configuration error turned on internet access for the three AI model instances that carried out the cyberattacks.

The most severe breach involved Claude Opus 4.7, an LLM that Anthropic released in April. The simulated company that it was asked to hack shared a name with a real website. The model subsequently hacked the organization that operates the website by chaining together multiple vulnerabilities.

Opus 4.7 compromised a production database with several hundred rows of information. Additionally, it obtained access credentials for several applications and infrastructure assets.

The second cyberattack was carried out by Mythos 5, Anthropic’s most advanced commercially available model. The LLM wrote a malicious Python package, or code bundle, and uploaded it to a popular open-source project hosting platform. The file was downloaded by a cybersecurity company a few minutes later. The malicious package compromised the firm’s infrastructure and stole access credentials.

According to Anthropic, the third cybersecurity incident involved an unnamed “internal research test model.” It compromised an application using a set of simple hacking methods such as SQL injections. At one point, the model discovered that the application wasn’t a part of its security evaluation sandbox and stopped the cyberattack.

Anthropic is partnering with a nonprofit AI safety lab called METR to carry out a more detailed investigation of the breaches. Additionally, the company plans to improve how it develops and monitors its LLM evaluation sandboxes. 

Image: Anthropic

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.

Sign in or create an account

SIGN IN

OR

New User? SIGN UP

Join us

SIGN UP

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry