Skip to content

OpenAI’s GPT-Red automates prompt injection testing to harden AI agents

8.2 relevance
Score Breakdown
technical depth
9
novelty
9
actionability
8
community
3
strategic
8
personal
10

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

OpenAI's automated prompt injection testing for AI agents, perfectly matches reader's interests.

AI/ML thenewstack.io
OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
Summary

OpenAI's GPT-Red automates prompt injection testing via self-play reinforcement learning, where an attacker model brute-forces exploit variations against defender models in simulated environments (emails, APIs, files). It successfully attacked nearly every evaluated model, and its findings helped harden GPT-5.6, which now shows a 0.05% failure rate on direct prompt-injection attempts—a 6x improvement over the previous strongest model. In live tests, GPT-Red manipulated an AI vending machine to drop prices to $0.50 and exfiltrated data from a Codex CLI agent using fewer tokens than a general-purpose frontier model.

Author

Amanda Caswell

More from Amanda Caswell →