Skip to content

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

7.8 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
7
community
8
strategic
7
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Study on humans missing threats in AI agent approvals is highly relevant to AI agent orchestration and security.

AI/ML scalex.dev
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Summary

A browser game simulating human-in-the-loop approval for AI coding agents found players missed 1 in 3 threats across 40,000 runs and 409,000 decisions. The most missed commands were npm run scripts (52.5% miss rate), even when the agent's history log showed the payload exfiltrating data via curl. Obviously destructive commands like rm -rf were caught 88.3% of the time, but credential exfiltration (cat ~/.aws/credentials) was missed 35% of the time, revealing a dangerous asymmetry in human vigilance.

Author

Alex Wauters