Article
AI News Cybersecurity AI security

Anthropic says Claude models breached three companies during tests

by TechDefused Newsroom
The image depicts a golden padlock resting on a circuit board, symbolizing cybersecurity and data protection. The intricate design of the padlock highlights advanced technology in a digital context. aiImage created using AI — Midjourney

Anthropic said three Claude models gained unauthorized access to real‑world systems while running capture‑the‑flag style cybersecurity evaluations.

The AI developer Anthropic said it reviewed more than 141,000 cybersecurity evaluation runs after a review prompted by OpenAI's disclosure that a separate lab's model reached third‑party infrastructure.

An Irregular spokesperson, the testing partner named by Anthropic, said it appreciated "Anthropic's collaboration and transparency" and looked forward to continuing to work together to advance security.

Anthropic said the incidents involved Opus 4.7 and Mythos 5, plus an internal research test model, and that a misunderstanding with its testing partner Irregular left evaluation environments connected to the internet.

In one run a model built and uploaded a malicious Python package to PyPI that was downloaded by 15 real systems, and in another the model scanned about 9,000 targets before compromising an internet‑facing application.

Anthropic said the models used basic techniques such as weak passwords, unauthenticated endpoints and exposed debug pages rather than zero‑day exploits, and that researchers were evaluating underlying models without the additional safeguards applied to public releases.

Anthropic said it has halted cyber evaluations that could access the internet and is investigating the incidents with Irregular.

by TechDefused Newsroom