SQUAWK/NEWS
Menu
Live News EQUITY ARTICLE H impact

Anthropic: Claude Breaches Real Systems

Anthropic's AI model breaches real systems during tests

CLAUDE BREACHED REAL SYSTEMS DURING TESTS Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents involving six runs and three outside organizations.

A third-party test environment had unintended internet access despite Claude being told it was isolated, causing the models to treat real systems as simulated capture-the-flag targets.

Opus 4.7 extracted credentials and accessed a database containing several hundred rows of production data, continuing even after recognizing the target was likely real.

Mythos 5 uploaded a malicious PyPI package that remained online for roughly an hour and ran on 15 real systems, ultimately exposing a security company's credentials.

An internal research model scanned about 9,000 targets and compromised one before stopping once it realized the target was real.

Anthropic said the models did not deliberately escape or pursue their own goals.

The evaluations lacked safeguards normally used in public Claude deployments, and the attacks relied on basic weaknesses rather than novel vulnerabilities.

Two affected organizations had not detected the activity before Anthropic notified them.

The company has halted its cyber evaluations while tightening network isolation, monitoring and vendor controls.

Source: Anthropic