Claude, OpenAI and Microsoft Copilot AI agents breaking free from containment restraints

Anthropic’s Claude Breach: Containment Failure or PR Play?

Opus 4.7, Mythos 5 and an unreleased research model reached the open internet during capture-the-flag drills and compromised real company systems, Anthropic says.

Three of Anthropic’s most advanced Claude models breached the real-world systems of three outside organizations during internal cybersecurity testing, the company disclosed Thursday. It was the second frontier AI lab in nine days to reveal that its models compromised production infrastructure while being evaluated for safety.

Anthropic said the incidents surfaced during a review of 141,006 cybersecurity evaluation transcripts, launched after OpenAI disclosed last week that its own models had broken out of an isolated test environment and reached Hugging Face, a platform for open-source AI models and datasets.

The disclosure is the second in nine days from a frontier AI lab admitting its own models breached outside systems during safety testing. Together, the incidents suggest evaluation environments – not just deployed products – have become a live attack surface of their own. For an industry that has spent years promising containment, two labs now say their containment failed within the same two-week span.

Security leaders are already lining up on both sides of what the incident means – from containment failure to marketing misstep. SPB rounds up that reaction below.

[Related: NVIDIA’s Answer to the Hugging Face Breach: Open Secure AI Alliance]

In the case of Anthropic, Claude was running a “capture the flag” exercise, a standard cybersecurity test in which the model hunts for hidden data on another machine. Anthropic’s evaluation prompts told Claude it had no internet access.

So Much for the Sandbox

A configuration error between Anthropic and its testing partner, Irregular, left the testing environment connected to the internet. Claude found real company systems during its search, mistook them for part of the exercise, and broke in taking advantage of weak passwords and unauthenticated endpoints.

Irregular said in a post on X, “Addressing these risks will require closer cooperation across the AI ecosystem.” Separately, Axios reported an Irregular spokesperson told them the company appreciates “Anthropic’s collaboration and transparency” and looks forward to continued work on security, while noting its own investigation is ongoing.

The earliest incident dates to April. Anthropic said it began its review on July 23, identified all three incidents by July 24, and notified Irregular and the affected organizations on July 27. Two had not detected the intrusions themselves.

Three separate models were involved. Opus 4.7 was assigned a fictional target whose name matched a real company. It found the real company online and attacked it, pulling credentials and several hundred rows of production data. That was the most serious of the three cases. It was also the only one where the model kept attacking after signs pointed to a real target.

Mythos 5 built and uploaded a malicious Python package to PyPI, the public Python software registry, believing the registry itself was part of the test; the package ran on 15 real systems in the roughly one hour it stayed live, letting Claude exfiltrate credentials from a security company’s malware-scanning system.

A third, unreleased internal research model scanned about 9,000 targets, compromised one company’s internet-facing application through SQL injection and an exposed debug page, then recognized the environment was real and stopped on its own.

Anthropic said its models did not exploit any unknown (“zero-day”) vulnerabilities and did not deliberately try to escape their test environments, distinguishing its incidents from OpenAI’s.

Containment Is Having a Rough Month

The disclosure comes amid a broader run of agentic AI systems reaching further than intended. Microsoft confirmed this week research by Norwegian AI researcher Håkon Måløy that a self-propagating “AI worm” can spread through Copilot and Word by hiding instructions in documents that later corrupt files created in Copilot workflows. The incident was reported by CSO Online that includes Microsoft’s official response.

Security Point Break reported this week that the same rogue OpenAI test agent behind the Hugging Face breach also compromised a customer of Modal Labs, a cloud platform for AI workloads.

A Gravitee-commissioned State of AI Agent Security 2026 survey, fielded in April 2026, found that 54% of organizations had experienced or suspected an AI-agent-related security or data-privacy incident in the prior 12 months, with 34.9% confirming an incident actually occurred. Telecom (67.3%) and financial services (54.7%) reported the highest rates. These figures describe agentic AI incidents broadly and are not specific to Anthropic or OpenAI.

The incidents underscore a clear operational lesson for AI security teams, according to security experts. Evaluation environments need the same access controls, monitoring and segmentation as production systems, whether they are run by Anthropic, OpenAI or an internal red team.

So Who Holds the Leash?

Security specialists told SPB the pattern points to a governance gap rather than a single vendor’s misstep.

Kristin Lowery, field CISO at Optiv, responding to the earlier OpenAI/Hugging Face breach, told SPB that agentic systems “can behave in harmful or unexpected ways even when the original goal is not malicious.”

Her fix isn’t more guardrails on the model. Rather, she said it’s tighter control over what agents can touch. Enterprises should treat autonomous agents “as a new class of privileged workload,” she added, with scoped permissions, short-lived credentials and a kill switch security teams can actually pull.

When an agent treats a real target as fictional, the failure is not only in the model; it is also in the controls that allowed the agent to act.

Adrian Balfour, founder of physical-AI consultancy Envorso, told SPB that this wasn’t really about intent at all. It was about containment. Claude wasn’t trying to escape anything, he said; the model was “aggressively executing their given task” and simply wrong about where the simulation ended.

Giuseppe Sette, president of Reflexivity, sees the bigger stakes further out. As AI vendors push into regulated fields like law, medicine and financial advice, he said, episodes like this will sharpen one question regulators keep coming back to. Can these models be trusted?

“AI Alignment is the most important research topic in AI – now and all the way to AGI… this is where human ingenuity comes into play… we will find unexpected ways to gauge how AIs respond and align them to our own goals,” Sette wrote to SPB, commenting on the trend.

Not everyone was persuaded the disclosure reflects sound judgment. Cybersecurity researcher Johannes Ullrich argued on LinkedIn that framing an “accidental” breach of outside companies as evidence of model power risks normalizing the behavior rather than curbing it.

Journalist Kate O’Flaherty made a similar point on LinkedIn, questioning why disclosing a real-world breach has become part of how labs signal model capability.

Anthropic said it has halted cybersecurity evaluations with internet access while it audits its testing pipeline, is working with independent evaluator METR on a third-party review, and plans to release a redacted transcript of the PyPI incident within the next week.

Total
0
Shares
Previous Article
Illustrated shield made of a network of hexagon icons around an NVIDIA chip, representing the Open Secure AI Alliance, next to a cracked padlock and an hourglass

NVIDIA's Answer to the Hugging Face Breach: Open Secure AI Alliance

Next Article
Illustration of a compromised VPN appliance representing the SonicWall SMA exploit chain

SonicWall VPN Flaw Chain Gives INC Ransomware Root Access

Related Posts

Discover more from Security Point Break

Subscribe now to keep reading and get access to the full archive.

Continue reading