Anthropic said on July 30 that it discovered instances in which its artificial intelligence model Claude gained “unauthorized access” to the systems of three organizations during cybersecurity testing.
The company said in a blog post that it made the discovery after reviewing 141,006 evaluation runs in which Claude could have obtained internet access.
Anthropic said that in the three identified incidents, Claude accessed the internet from within or while interacting with the evaluation environment of one of its third-party evaluation partners.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” it stated.
“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
The review was launched after OpenAI disclosed on July 21 that several of its models had broken out of an isolated test environment and accessed the production infrastructure of Hugging Face, a platform for open-source machine learning models and AI datasets.
This story will be updated with additional details.





















