Prompt Injection in Claude Code Opus 5 Auto Mode
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Prompt injection in Claude Code Opus 5, directly relevant to AI agent security and orchestration.
A targeted attack chain achieves 60-80% prompt injection success against Claude Code Opus 5 in Auto Mode, contradicting Anthropic's commissioned third-party evaluation that reported 0.00% success. The attack works by nudging Claude from the WebFetch tool to curl, redirecting to a ZIP archive containing a malicious struct.py that shadows Python's standard library, triggering code execution when the model imports base64. Auto Mode, now default since mid-August, replaces human approval with a safety classifier but does not substitute for running agents in isolated environments.
wunderwuzzi