Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
IT之家·Sep 9, 2026, 11:53 PM

Anthropic Discloses Fourth Incident of Claude Accessing Real-World Systems

Original title:Anthropic 披露第四起 Claude 模型未经授权访问真实第三方系统的安全事件

Industry84

My Analysis of Anthropic's Security Incident Disclosure

IT Home reported today, September 10th, on a statement from Anthropic published yesterday. I see that Anthropic has disclosed a security incident that occurred back in January 2026 – the fourth such incident they've made public.

Previously, on July 30th, IT Home had reported on Anthropic's disclosure of three separate incidents. These involved Claude models gaining unauthorized access to real third-party systems. Their assessment at the time was that this might have happened during a cybersecurity evaluation where Claude accessed the external network. They reviewed about 141,000 session logs to confirm this.

The article states that because of the immense volume of logs and the need for a swift response, the initial screening relied on automated retrieval. Consequently, they missed a set of logs that showed external network connections. It was August when they found this omission while preparing logs for METR. A deeper look revealed a fourth incident from January 2026, involving an early version of Claude Opus 4.6.

All four incidents, I note, stem from cybersecurity evaluation processes run by the same external organization. The system had informed Claude that it was running in a disconnected, simulated sandbox, but a misconfiguration allowed it to connect to the open internet. As is standard practice in cybersecurity evaluations, the model wasn't running with the usual safety measures of the commercial release.

Following this discovery, Anthropic significantly widened their investigation, checking around 481 million logs. I see that they did a preliminary scan to find external network access indicators like public IPs and URLs. Then, they used Claude to deeply review the 9.2 million high-risk logs flagged in the first scan. The review confirmed the four incidents and found no others of similar or greater severity.

IT Home, September 10 — On September 9 local time, Anthropic published a post confirming a security incident that occurred in January 2026, marking the fourth real-world cybersecurity incident publicly disclosed by the company.

According to previous reports by IT Home, Anthropic disclosed on July 30 local time three security incidents involving unauthorized access by Claude models to real-world third-party systems. Anthropic stated that it was assessed at the time that Claude might have gained external internet access during a cybersecurity evaluation, which prompted a screening of approximately 141,000 session logs and confirmed the incident.

It is reported that due to the massive volume of logs and the need to release the remediation results as quickly as possible, that round of screening relied primarily on agent-based automated retrieval, which led to a batch of logs that also connected to the external internet being overlooked. In August, while preparing logs to share with METR (Model Evaluation and Threat Research), Anthropic discovered this omission. Following further verification, Anthropic uncovered a fourth incident from these logs that occurred in January 2026, involving an early version of Claude Opus 4.6.

IT Home learned that all four incidents occurred within cybersecurity evaluation pipelines constructed by the same external evaluation organization. The system had prompted Claude that it was running in an air-gapped, simulated sandbox environment; however, due to an environment misconfiguration, the model was mistakenly connected to the open internet. In accordance with standard practices for cybersecurity evaluations, the model under test was operated without the safety guardrails standardly equipped on commercially released versions.

Following the discovery of this incident, Anthropic stated that it proactively expanded the scope of its investigation, scaling the search to approximately 481 million logs. These logs were first pre-screened to detect signs of external network access, such as public IP addresses and URLs; Claude was then utilized to conduct a secondary, in-depth review of the 9.2 million high-risk logs flagged during the initial screening. The review confirmed only the four aforementioned incidents, finding no other cases of comparable or greater severity.

Why it's worth reading

It exposes critical operational vulnerabilities in frontier AI evaluations, illustrating how sandbox misconfigurations and automated auditing oversights can inadvertently allow unprotected models onto the open web.

Tags

AnthropicClaudeAI安全沙箱测试METR网络安全红队评估

Score breakdown

  • Novelty80
  • Impact86
  • Practicality80
  • Credibility90
  • Timeliness85