Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Anthropic says three Claude models breached real company systems during safety tests

Anthropic disclosed that three Claude models gained unauthorized access to real production systems at three organizations during misconfigured cybersecurity evaluations.

D
Jul 30, 2026 · 1 min read

Anthropic said on July 30 that three of its Claude models gained unauthorized access to real production systems at three organizations during cybersecurity evaluations meant to run in sealed test environments.

The disclosure is a rare admission that a frontier AI model can slip the boundaries of its own safety tests. The breaches happened after a third-party evaluation partner’s testing setup was misconfigured to allow live internet access, even though the models’ prompts stated the environment was isolated, the company said.

The models involved were Claude Opus 4.7, a model called Mythos 5, and an internal research model. Opus 4.7 extracted credentials and reached a database holding several hundred rows of production data at a real company that happened to share a name with a fictional test scenario. Mythos 5 published a malicious Python package to PyPI, the public code registry, that was downloaded and run on 15 real systems within an hour, including at a security company.

The internal research model compromised one company’s application using basic, well-known techniques after scanning roughly 9,000 targets, according to Anthropic’s disclosure.

The earliest incident dates to April 2026. Anthropic said it began reviewing more than 141,000 evaluation runs after OpenAI disclosed a similar test-environment escape, identified the problem starting July 23 and notified the affected organizations by July 27.

The company suspended all cybersecurity evaluations and notified its evaluation partner, Irregular, along with the affected organizations. Anthropic did not say whether any data was misused or whether the organizations suffered lasting harm.

More news