Claude AI Escaped Test Sandbox to Attack Three Organizations
Wrote and published malware during tests, which is apparently OK because leaky test environments were the real problem
Menu
Front Page Breaking News Comments Flagged Comments Recently Flagged User Blogs Write a Blog Entry Create a Poll Edit Account Weekly Digest Stats Page RSS Feed Back Page
Subscriptions
Read the Retort using RSS.
RSS Feed
Author Info
lamplighter
Joined 2013/04/13Visited 2026/08/25
Status: user
MORE STORIES
US Cancels Marine Drills with South Korea (2 comments) ...
Florida feels the squeeze of rising healthcare costs (2 comments) ...
DHS Confirms Just 185 Non-citizens on Nevada Voter Rolls (15 comments) ...
US Says It Aided Passage of Oil Through Strait of Hormuz (14 comments) ...
Sullivan Endorses Sullivan Against Sullivan (6 comments) ...
Alternate links: Google News | Twitter
Anthropic's AI Claude escaped testing environment and hacked organizations[image or embed] -- The Guardian (@theguardian.com) 8:15 PM · Jul 30, 2026
Anthropic's AI Claude escaped testing environment and hacked organizations[image or embed]
Admin's note: Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.
More from the article ...
... "In particular, we looked for evidence that Claude -- like the OpenAI models that accessed Hugging Face --was able to access the internet from within testing environments that should have been sealed off," Anthropic wrote. The company considered 141,006 evaluation runs during which Claude could have obtained internet access and found "three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations." Anthropic's code made those intrusions while participating in capture-the-flag challenges, tests that challenge attackers to retrieve a piece of information. Human hackers often participate in capture-the-flag tests, so figuring out how AI tackles such tasks is of interest. Anthropic works with a company called Irregular to conduct tests of this sort. ...
The company considered 141,006 evaluation runs during which Claude could have obtained internet access and found "three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations."
Anthropic's code made those intrusions while participating in capture-the-flag challenges, tests that challenge attackers to retrieve a piece of information. Human hackers often participate in capture-the-flag tests, so figuring out how AI tackles such tasks is of interest. Anthropic works with a company called Irregular to conduct tests of this sort. ...
#1 | Posted by LampLighter at 2026-07-31 01:47 PM | Reply
Seems Claude is an arse.
Probably Antifa, too.
#2 | Posted by donnerboy at 2026-08-02 08:47 PM | Reply
Humans have ------ up the world enough.
Time to let the computers finish us off.
#3 | Posted by Nixon at 2026-08-03 12:07 PM | Reply
AI is going through people's voice, email and text communications and using it to blackmail them with it. I forget who it was, but they were saying it's up to lick 65% of AI is doing this in test runs.
I hope JPW doesn't book his appointments through texts or something like it.
#4 | Posted by lfthndthrds at 2026-08-03 05:23 PM | Reply
XV
#5 | Posted by lfthndthrds at 2026-08-03 10:12 PM | Reply
Post a commentComments are closed for this entry.Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy
Comments are closed for this entry.
Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy