Frozen copy retrieved 2026-08-01T18:22:11Z for audit 2026-08-01T19-18-56Z. Original URL: https://1stskyomaha.com/anthropic-says-its-ai-models-hacked-3-organizations-during-testing/. The Stochastic Parrot does not host or redistribute; this snapshot exists solely so that quoted spans remain verifiable if the original page changes. Character offsets below index into this plain text; highlighted spans are the quotes cited in the audit.

Anthropic says its AI models hacked 3 organizations during testing

Associated Press (via 1st Sky Omaha) · back to the audit
Anthropic said its artificial intelligence models hacked into three other organizations during testing.

The disclosure came just days after ChatGPT maker OpenAI raised concerns over AI controls when it disclosed its own rogue models had hacked another company.

Anthropic, the San Francisco-based company behind Claude, posted Thursday on its website that it found the three incidents after reviewing more than 141,000 evaluation runs. It called the effort a "large-scale" cybersecurity review, launched in response to the OpenAI incident, to check whether its models could reach the internet from testing environments that should have been sealed off.

The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model, with the earliest incidents dating to April. In each case, the model was given a "capture the flag" challenge - a fictional scenario where a secret piece of information (the "flag") was hidden on another machine on the network, with the goal of breaking in to retrieve it.

Anthropic said "Claude compromised the impacted organizations' infrastructure using basic techniques," such as exploiting weak passwords. It did not name the three organizations. Two said they had not previously detected the activity, and Anthropic said it was "continuing to reach out to the third."

The review was done with Irregular, which describes itself as the "first frontier security lab." Irregular said in a post on X: "Addressing these risks will require closer cooperation across the AI ecosystem."

Last week, OpenAI said its models went rogue during an evaluation, breaking into the servers of AI startup Hugging Face - an incident OpenAI described as a "significant security incident."

Anthropic said on its website: "Safety testing happens before a model is released precisely because we don't yet know what it is capable of."