Security products are routinely evaluated on true positive rates alone. False positive rates are rarely published, rarely measured with statistical rigor, and almost never independently verified. In AI agent behavioral anomaly detection, a single false positive that blocks a production agent can terminate a pilot deployment permanently. This paper proposes a complete false positive rate measurement methodology including: a specification-first corpus construction protocol eliminating co-design bias, a 9-category behavioral taxonomy covering the full range of legitimate enterprise agent archetypes, Clopper-Pearson exact confidence intervals for zero-FP observations, a production measurement protocol using customer-verifiable SQL, and an auto-rollback circuit breaker as a compliance control. Implemented in AgentRepEngine (DOI: 10.5281/zenodo.19169185).
Rehan Masood (Thu,) studied this question.