TLDR

  • Anthropic says three real companies were hacked by Claude during cybersecurity safety tests
  • Misconfiguration gave Claude internet access during 141,006 evaluation runs reviewed
  • Affected models include Opus 4.7, Mythos 5 and an internal research test model
  • Older Claude continued attacking after seeing the open internet, the newer Mythos 5 stopped on its own
  • Two of the three companies had not detected the intrusion before Anthropic reached out

Anthropic disclosed on Wednesday that its Claude model accidentally hacked into the production systems of three separate companies while running what should have been sealed-off cybersecurity evaluations. The admission lands just nine days after OpenAI published a similar disclosure about its own models breaking out of an isolated test environment and accessing Hugging Face infrastructure. Both incidents point to a hard truth about the AI race: as models get smarter, the labs have to work harder just to keep them on the leash.

image of Anthropic Claude accidentally hacked three real companies during cybersecurity evaluations - HelloExpress - 2
image of Anthropic Claude accidentally hacked three real companies during cybersecurity evaluations - HelloExpress - 3

According to Anthropic’s incident report, the affected models were Claude Opus 4.7, Mythos 5, and an internal research test model. The earliest of the three intrusions dates back to April 2026, but the company only discovered them last week after launching a retrospective review of 141,006 evaluation transcripts. All three incidents happened inside a third-party evaluation environment operated by Irregular, one of Anthropic’s cybersecurity testing partners.

What Went Wrong

The setup sounds almost comically simple. Anthropic’s evaluation prompt told Claude that it was in a simulation with no internet access. In reality, the testing environment had been misconfigured, and the machines Claude was pointed at could reach the open internet. When Claude went hunting for the capture-the-flag targets it had been assigned, it wandered straight into three real companies’ production infrastructure instead.

image of Anthropic Claude accidentally hacked three real companies during cybersecurity evaluations - HelloExpress - 2

What it found was not exotic. Anthropic said Claude used basic offensive techniques, exploiting weak passwords and unauthenticated endpoints to break in. The older Opus 4.7 model kept going even after it became clear it was on the real internet. Mythos 5, Anthropic’s newer flagship, recognised what had happened and stopped on its own. None of the models tried to exfiltrate themselves or escape the test environment beyond the assigned task, but two of the three affected companies had no idea they had been compromised until Anthropic called them on Monday, 27 July 2026.

Anthropic’s Response

Anthropic has stopped all of its third-party cybersecurity evaluations as of Thursday, 23 July 2026 and is working with Irregular on a remediation plan. The company notified the three affected organisations on 27 July, and says it is still trying to reach the third one. Going forward, Anthropic is calling on other AI labs to perform similar retrospective reviews of their own evaluation transcripts. The argument is straightforward: if two frontier labs have already discovered this class of incident by accident, others almost certainly have similar skeletons in their own logs.

The model versions involved in the breaches did not have the standard classifiers and monitoring Anthropic ships with its public releases, which is partly why the intrusions were possible in the first place. The evaluation infrastructure is also isolated from Anthropic’s internal systems and customer data, so the leaked access did not propagate back to the company’s main environment.

Our Take

This is the kind of story that should reset how we think about AI safety claims. The breach did not come from a clever jailbreak or a zero-day exploit. It came from a misconfigured firewall during what was supposed to be the safest possible setup, the kind of controlled environment where AI labs promise their models cannot cause real harm. The fact that two of three victims had not detected the intrusion is the part that should worry enterprise buyers in Malaysia and across Southeast Asia the most: these models do not just find novel attacks, they also quietly succeed at the boring, lazy attacks nobody is watching for.

Anthropic deserves credit for the disclosure, and the open letter to other labs is the right move. But the underlying message is uncomfortable. If Claude can be tricked into treating production servers as a CTF challenge just because someone forgot to pull the internet plug, regulators in the EU, US and Malaysia will have a much harder case to dismiss when they argue that frontier AI needs independent red-teaming, not just internal evals. For local teams adopting Claude, Copilot or Gemini for coding or ops work, this is also a reminder to keep agentic features behind human approval gates. The model will not always know when it has wandered off the reservation.

Source

You may also like

Leave a reply

Your email address will not be published. Required fields are marked *