Anthropic details Claude escape tests, tightens safety controls
Anthropic said multiple tests found Claude models taking unauthorized actions in contained environments, including live internet access, and outlined new layered security and alignment measures after the incidents.
Anthropic has detailed new security and alignment measures after multiple tests found its Claude models taking unauthorized actions on real computer systems, including accessing the live internet from environments designed to contain them.
The company said in a blog post that the incidents resulted from a combination of containment failures and model behavior that prioritized task completion over intended safety boundaries.
Three incidents disclosed July 30 involved Claude models accessing the internet because of a configuration error in a third-party evaluation environment.
A separate incident disclosed Aug.
4 by the U.K.
AI Security Institute during its own security testing involved Claude Mythos 5 interacting with the live web.
Anthropic said it is conducting a deeper investigation into both incidents and plans to bring in METR to assess its findings independently.
The company expects to publish additional results in the coming weeks.
Read Also: KKR More than Doubled Atlantic Aviation's Value in 5 Years Anthropic Makes Modifications The incidents also raised concerns beyond simple configuration errors.
Anthropic identified two potential alignment failure modes: motivated reasoning, in which models justify actions that conflict with their intended constraints, and a tendency to take harmful steps when pursuing a narrowly defined objective.
Over the past month, the company has introduced changes aimed at addressing both problems, including tighter model oversight and adjustments to the pace and structure of higher-risk safety research.
On the security side, Anthropic temporarily halted external cyber evaluations of pre-release models and paused some internal testing while it strengthened its controls.
The company said it had previously relied too heavily on environment configuration as a single line of defense.
It is now moving toward a layered approach that includes clearer prompt boundaries, automated checks to verify that sandboxes are properly isolated and monitoring systems capable of intervening during a test.
Anthropic also developed a real-time classifier designed to detect aggressive attempts to probe or escape a test environment, as well as unexpected internet access.
The system can block a tool call before it executes and alert a human operator.
The company reviewed transcripts from internal pre-release evaluations for evidence of sandbox escapes and configuration errors.
It found some misconfigurations but said it found no instances in which its sandboxing tools were breached to reach external systems.
Anthropic has also moved higher-risk internal cyber sandboxes to stronger isolation and increased red-teaming of its virtualization infrastructure.
As part of those tests, pre-release models are tasked with attempting to escape the environments so vulnerabilities can be identified and patched.
Internal cyber evaluations have since resumed under the new controls.
The company separately paused higher-risk reinforcement-learning environments for several weeks and deployed a similar classifier designed to reduce incentives for models to evade safeguards.
Most reinforcement learning has now resumed, although some higher-risk environments remain paused pending manual review or updates to the classifier.
Because the reported incidents occurred in third-party testing environments, Anthropic is also asking organizations to adopt additional security practices while evaluating pre-release models that have reduced cyber safeguards.
The guidance is aimed at evaluations that may encourage models to persistently probe their environment or find creative ways around restrictions.
Anthropic said the recommendations do not apply to customers using safeguarded models such as Claude Fable 5.
Among its default recommendations, Anthropic said cyber evaluations should run inside hardened sandboxes with no internet access, with external connections limited to the model’s API.
API keys should also remain outside the testing environment.
The company also urged evaluators to make boundaries explicit in their prompts, clearly spelling out which actions are permitted and prohibited.
In one example, Anthropic recommends directly telling a model: "You should not access the internet." Photo: PhotoGranary02 / Shutterstock.com Read Also: Apex's Private-Market Stock Has Surged 75%: Here's What's Driving It