This research evaluates accuracy and risk in deep learning systems, suggesting improved reliability for safety-critical applications.
Despite achieving excellent performance on benchmarks, deep neural networks often underperform in real-world scenarios due to their sensitivity to minor shifts in input data, referred to as distributional shifts. These shifts are common in practical scenarios but are rarely accounted for during evaluations, leading to inflated performance metrics. To address this gap, we propose a novel methodology for the verification, evaluation, and risk assessment of deep learning systems. Our approach explicitly models the incidence of distributional shifts at runtime by estimating their probability from the outputs of out-of-distribution detectors. We combine these estimates with conditional probabilities of network correctness, structuring them in a binary tree. By traversing this tree, we can compute reliable and precise estimates of network accuracy. We assess our approach on five datasets, simulating deployment conditions characterized by different frequencies of distributional shift. Our approach consistently outperforms conventional evaluations, with accuracy estimation errors typically ranging between 0.01 and 0.10. We further showcase the potential of our approach on a medical segmentation benchmark, wherein we apply our methods to risk assessment by associating costs with tree nodes, informing cost-benefit analyses and decision-making. Overall, our approach offers a robust framework for improving the reliability and trustworthiness of deep learning systems, particularly in safety-critical applications, by providing more accurate evaluations and actionable risk assessments.
No takes yet. Share an insight, caveat, or question.
Torpmann-Hagen et al. (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: