Skip to content

I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

7.9 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
9
community
6
strategic
5
personal
10

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Field test of 170 agent goals for $0.49 uncovering issues unit tests miss, highly actionable for agent testing and observability.

AI/ML dev.to
I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.
Summary

A field test of the PlannerCritic agent engine across 170 goals at $0.49 found zero issues in v0.2.1, but only because code review caught all 41 bugs before the LLM ran. In v0.1.0, the same test found 10 issues—including harness bugs, prompt gaps, and design flaws—that unit tests missed entirely, such as 57 of 65 assertion files being in the wrong format with no error. The arc shows the field test evolving from a diagnostic tool into an immutable regression gate, where zero failures now signals rigorous pre-release review rather than a broken harness.

Author

Debashish Ghosal

More from Debashish Ghosal →