Skip to content
Frozen copy retrieved 2026-08-01T18:22:11Z for audit 2026-08-01T19-18-56Z. Original URL:
https://community.triblive.com/news/4105600. The Stochastic Parrot does not host or redistribute; this snapshot exists solely so that quoted spans remain verifiable if the original page changes. Character offsets below index into this plain text; highlighted spans are the quotes cited in the audit.
Anthropic's AI models hacked 3 organizations during tests
Anthropic PBC announced that its artificial intelligence models had breached three different organizations during cybersecurity tests that went awry a little more than a week after its chief rival, OpenAI, disclosed a similar incident.
Anthropic said in a Thursday blog post that it uncovered the problem by reviewing its own cybersecurity evaluations, after OpenAI disclosed a similar event. In both cases, "the AI models were able to access the internet from within testing environments that should have been sealed off."
The company said it "reviewed 141,006 evaluation tests" and "found three instances in which its Claude AI tool accessed the internet and then hacked into" "the real-world infrastructure of external organizations," with "the earliest incidents date to April."
"The affected organizations aren't named in the blog." They "don't include Hugging Face or the cloud platform Modal," "according to a person familiar with the matter." "The OpenAI models that hacked Hugging Face also gained access to a Modal customer's account."
"Neither Anthropic nor the organizations that were breached had noticed the intrusions." Anthropic acknowledged it "could have done more to review network logs and evaluation transcripts."
According to the blog, the breaches involved "Opus 4.7, Mythos 5 and an internal research test model," and "Each of the models ran without the safeguards normally implemented in public tools." "Claude compromised the organizations using basic techniques such as exploiting weak passwords, according to the blog."
All three incidents happened in "evaluation environments built by the AI security firm Irregular." "In every case, Anthropic specified to Claude that its environment was a simulation and that it had no internet access." "That was wrong, according to Anthropic." The blog explains: "Due to a misunderstanding between us and our evaluation partner, this was not the case." "When the models gained unauthorized access to three other organizations, it treated each of them as part of an exercise." "The tests were 'capture-the-flag' evaluations" - "a common way to test hacking capabilities for both humans and AI."
"An Irregular spokesperson said the company appreciates Anthropic's collaboration and transparency." "The company's investigation is ongoing, the spokesperson said."
"More than 1,100 staffers across artificial intelligence firms also signed a petition on Tuesday," "as Bloomberg first reported," urging the U.S. government to support a mechanism that would help "deliberately pace" AI development.
The disclosure came three months after the company announced "a new AI model known as Mythos" that was "so powerful and potentially dangerous that the company strictly limited its release."