Intrusion detectors are often evaluated using average metrics at unconstrained thresholds, yet deployments require explicit control over false alarms. We investigate zero-day (out-of-distribution, OOD) intrusion detection under a target-FPR calibrated protocol, where a threshold is set on benign validation traffic to satisfy a target false positive rate α and transferred, unchanged, to a seen-test and OOD-test. Using CICIDS2017-derived host-session nodes aggregated in 1min and 5min windows, we compare tabular baselines, message-passing GNNs on a rule-based graph, and employ a method that builds a k-nearest-neighbor similarity graph with lightweight feature pre-smoothing. Robustness is measured using the OOD violation ratio, percentile tail risk, and feasibility under explicit false-alarm budgets. Base-graph GNNs exhibit heavy-tailed false-alarm amplification under OOD shifts: at α = 0.001, the p95 violation ratio reaches 68.50 (1m) and 67.95 (5m). In contrast, the proposed method reduces p95 to 3.41 (1m) and 1.15 (5m) and improves budget feasibility. We further verify robustness beyond a single held-out family by evaluating additional unseen-family splits (e.g., DDoS and DDoS+DoS) under the same calibrated operating point. We also quantify deployment-oriented cost via edge-list size and practical parsing/loading time. These findings suggest that similarity-based graphs with light pre-smoothing improve deployability under distribution shifts.
Ha et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: