This paper reports the detection performance of the TACET network defense platform under sustained adversary emulation. Three prior papers described the platform architecture, its hybrid post-quantum attestation protocol, and a 41-day passive live-network soak that established operational stability but generated no ground-truth attacks. That gap is the motivation here. Over 26 days, from June 13 to July 9, 2026, an autonomous red-team generator fired 2,414 labeled attacks across four parallel campaigns at a TACET v4.8 sensor watching a live /24 network of 141 tracked devices. Every fired attack wrote one ground-truth record, and recall was reconstructed by joining those records to the sensor's own output over an 11 GB archive of logs, a detections database, and per-detector event streams. Every headline figure was recomputed from primary data and cross-checked three independent ways. The result is bimodal. Detection of overt, discrete attacks is production grade. Nine of ten red-team families sat at or above 92 percent, five at 100 percent, and every attack the sensor itself rated CRITICAL was caught, 170 of 170. Excluding one deliberately stealthy family, red-team recall was 86.9 percent. The cross-layer trust engine independently drove the attacker's identity score to 0.544 against a 0.957 network mean, distinguishing the hostile host without being told which host was hostile. The two real coverage gaps are both low-observable by design: low-and-slow scanning at 37 percent per probe, though the campaign was correlated into a standing alert within about three minutes, and unicast command-and-control beacons at 14 percent. The dominant operational problem is precision, not recall. The behavioral attack-graph raised 41,189 chains, roughly 19,000 at threat score 0.9 or higher, while the attacker produced only 226 of them. The largest single false-positive source is the sensor's own address, whose neighbor-discovery traffic is read as scanning and drives 9,564 self-directed automated responses. At rest, over eight attack-free days, the sensor raised about 81 operator alerts per day against roughly two genuine detections, a signal-to-noise ratio near one in forty. We show that this floor is concentrated in a handful of devices, that raising confidence thresholds cannot fix it because the scores do not separate attacker from benign, and that the fix must be structural. We give six prioritized recommendations. Recall is production-grade on the evaluated attack set and network. The false-positive floor needs a tuning pass before the automated response layer can be trusted to act unattended. These results are bounded to the network layer, one live /24, and one non-adaptive attacker; Section 23 states the limitations in full.
Alexander W. Smith (Sun,) studied this question.