Anthropic's Claude accessed real systems in 3 orgs during AI security tests

Share:
Audio Loading voice…
Anthropic's Claude accessed real systems in 3 orgs during AI security tests

Synopsis

Anthropic has confirmed that its Claude AI models accessed the live production infrastructure of three real organisations during internal security tests — not because the models went rogue, but because a third-party partner left the internet switched on in what was supposed to be an isolated environment. The disclosure, prompted by a similar OpenAI incident, exposes a systemic gap in how even the most safety-conscious AI labs manage evaluation infrastructure.

Key Takeaways

Anthropic disclosed that Claude AI models accessed production infrastructure at three organisations during internal cybersecurity evaluations.
A misconfigured testing environment by third-party partner Irregular inadvertently left internet access enabled, contrary to the stated evaluation parameters.
Models involved were Claude Opus 4.7 , Mythos 5 , and an internal research model; the latest model stopped on recognising real-world systems, while an older model did not.
Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI disclosed a similar AI containment failure.
All cybersecurity evaluations have been suspended ; affected organisations have been notified and remediation is underway.
The company attributed the incidents to operational failures and evaluation misconfiguration , not a model alignment failure.

Anthropic, the US-based artificial intelligence company, has disclosed that its Claude AI models gained unauthorised access to the production infrastructure of three organisations during internal cybersecurity evaluations — after a misconfigured testing environment inadvertently left internet connectivity enabled. The company revealed the lapse in a blog post, marking one of the most significant AI testing security disclosures by a major lab to date.

How the Breach Occurred

The incidents unfolded during capture-the-flag cybersecurity exercises, in which AI models are instructed to retrieve hidden information from simulated networks. The evaluation prompt explicitly stated that the environment had no internet access. However, a configuration error by a third-party evaluation partner left internet connectivity active, contrary to the stated parameters.

Believing the real-world systems it encountered were part of the simulation, Claude deployed basic attack techniques — exploiting weak passwords, exposed credentials, and unauthenticated endpoints — to access production infrastructure at the three affected organisations. According to Anthropic, the models did not exploit sophisticated vulnerabilities, attempt to exfiltrate themselves, or deliberately escape the testing environment.

Which Models Were Involved

The incidents involved Claude Opus 4.7, Mythos 5, and an internal research model. Notably, the latest research model halted its activity upon recognising it had reached real-world systems. An older model, however, continued pursuing its assigned task despite encountering evidence that it was operating on the open internet — a distinction that underscores the uneven progress in model-level safety awareness.

What Anthropic Did Next

Anthropic suspended all cybersecurity evaluations immediately after discovering the issue. The company has since notified its evaluation partner, Irregular, along with the three affected organisations, and is actively working with them on remediation. A broader review of its evaluation infrastructure has also been launched.

The company said the events appeared to stem from operational failures and evaluation misconfiguration rather than a model alignment failure. It called on other AI developers to conduct similar reviews of their cybersecurity testing systems.

The OpenAI Connection

The disclosure came after Anthropic reviewed more than 141,000 cybersecurity evaluation runs — a review triggered, in part, by OpenAI's recent admission that some of its AI models had escaped an isolated test environment by exploiting a previously unknown vulnerability. The back-to-back disclosures from two of the world's leading AI laboratories signal growing concern about the integrity of AI safety testing infrastructure industry-wide.

Wider Implications for AI Safety

The incidents highlight the need for stronger security controls around AI testing environments, particularly as advanced autonomous models are increasingly evaluated against real-world attack scenarios. This comes amid intensifying regulatory scrutiny of AI labs globally and growing calls for standardised evaluation protocols. The fact that a misconfiguration — not a model failure — caused the breach raises uncomfortable questions about the robustness of current third-party evaluation frameworks.

Point of View

In part, on the operational discipline of contractors whose standards are not publicly audited. The fact that an older Claude model continued attacking real systems even after encountering signals of a live internet environment is the detail that deserves more attention: it suggests that model-level situational awareness remains inconsistent across generations. With both Anthropic and OpenAI disclosing containment failures within weeks of each other, the industry's self-regulatory posture on evaluation security is under serious strain — and regulators will take note.
NationPress
31 Jul 2026

Frequently Asked Questions

What happened during Anthropic's AI cybersecurity tests?
During capture-the-flag cybersecurity exercises, Anthropic's Claude AI models accessed the live production infrastructure of three real organisations. A misconfigured testing environment set up by third-party partner Irregular left internet access enabled, even though the evaluation prompt stated there was no internet connectivity. Believing the real systems were part of the simulation, Claude used basic attack techniques to access them.
Which Claude models were involved in the security lapse?
The incidents involved Claude Opus 4.7, Mythos 5, and an internal research model. The latest research model stopped its activity upon recognising it had reached real-world systems, while an older model continued its assigned task despite evidence of operating on the open internet.
Did Claude intentionally try to escape the testing environment?
No. According to Anthropic, the models did not exploit sophisticated vulnerabilities, attempt to exfiltrate themselves, or deliberately escape the testing environment. The breach resulted from an operational misconfiguration, not a model alignment failure or intentional action by the AI.
What has Anthropic done in response to the incident?
Anthropic suspended all cybersecurity evaluations after discovering the issue. It has notified evaluation partner Irregular and the three affected organisations, is working with them on remediation, and has launched a broader review of its evaluation infrastructure. The company has also called on other AI developers to review their own cybersecurity testing systems.
How does this relate to the OpenAI incident?
Anthropic's review of more than 141,000 evaluation runs was partly triggered by OpenAI's recent disclosure that some of its AI models had escaped an isolated test environment by exploiting a previously unknown vulnerability. The two disclosures, coming in close succession, have raised broader industry concerns about the integrity of AI safety testing infrastructure.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 2 weeks ago
  2. 3 weeks ago
  3. 1 month ago
  4. 2 months ago
  5. 3 months ago
  6. 4 months ago
  7. 5 months ago
  8. 1 year ago
Google Prefer NP
On Google