Pre-deployment evaluation of language models suffers from an unresolved stopping problem. How much testing is enough to support a transparent and defensible release decision? Existing guidance emphasizes lifecycle test, evaluation, verification, and validation (TEVV), deployment-specific evidence, and post-deployment monitoring, but it does not provide a widely adopted quantitative stopping rule for evaluation sufficiency. This paper proposes a dual-lane sequential method. Lane A uses representative testing to bound failure rates within deployment-critical slices via one-sided exact binomial confidence limits. Lane B uses targeted or adversarial testing to assess whether the tail of novel severe error mechanisms is saturating, using discovery curves, rolling novelty, and Good-Turing missing-mass estimates. The resulting stopping rule advances a system only when severe-failure risk is bounded below an agreed threshold, and the targeted discovery process has flattened enough that additional testing is unlikely to materially change the release decision. The paper contributes three elements: (i) a formal problem statement for deployment readiness stopping, (ii) a practical composite rule that joins rate bounds and tail saturation, and (iii) a Monte Carlo study comparing the proposed rule against fixed-budget, recent-novelty, and rate-only baselines. In the safe long-tail simulation, premature stops fell from 25.0% under a rate-only rule and 100% under a fixed budget to 5.0% under the proposed dual rule, at the cost of additional evaluations. In an unsafe long-tail scenario, recent-novelty stopping still stopped 100% of the time, whereas the dual rule refused to stop within budget. A synthetic case study for a policy-oriented retrieval assistant illustrates how the method can be instantiated in practice. The proposed method does not prove safety or correctness. Its purpose is to provide an explicit, reviewable evidence structure for arguing that residual deployment risk has been reduced to an acceptable level for a given release stage.
Richard Heimann (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: