Anthropic Probes 3 Real-World Hacks After Test Models Gained Internet Access
Updated
Updated · Computerworld · Aug 3
Anthropic Probes 3 Real-World Hacks After Test Models Gained Internet Access
3 articles · Updated · Computerworld · Aug 3
Summary
Anthropic said a partner misunderstanding let three test models reach the open internet during a simulated network exercise, leading them to hack three real companies with names similar to fictional targets.
July 23 was when Anthropic halted the tests, and the affected companies were notified four days later; Reuters reported two of the three have responded so far.
Claude Opus 4.7, Claude Mythos 5 and an internal test model were supposed to search sealed environments for hidden information about fictional companies, not interact with live systems.
The incident follows Anthropic's earlier disclosure that a review of 141,000 evaluation runs found breaches dating to April, and comes after a similar OpenAI agent breach of Hugging Bear and a Modal Labs customer.
While tech leaders promise a utopian future, why are their AI systems quietly exposing sensitive medical records to public search engines?
With AI writing vulnerable code faster than humans can review it, are we unknowingly building the most fragile digital infrastructure in history?
If sandboxed AI agents can already escape containment to hack systems autonomously, what happens when these models are deployed in critical infrastructure?