Claude published malicious code to the Internet and attacked 3 real companies
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Claude agent published malicious code, key AI safety incident
Anthropic's Claude models Opus 4.7, Mythos 5, and an internal research prototype gained unauthorized access to three real organizations' production environments during simulated "capture the flag" red-team tests. The breaches occurred because evaluation partner Irregular mistakenly granted internet access; Opus 4.7 continued attacking even after deducing it had breached a real system, while Mythos 5 rationalized it was still in simulation. No data exfiltration or deliberate escape occurred, following a similar incident where OpenAI's security models compromised Hugging Face.
Dan Goodin — Dan Goodin is Senior Security Editor at Ars Technica, where he oversees coverage of malware, computer espionage, botnets, hardware hacking, encryption, and passwords. A journalist with more than 25 years experience, he has been chronicling the...