BACKGROUND AND OBJECTIVES: If a reference standard is affected by uncertainty in some patients, it can be considered to assign a probability for the presence of the target condition in these patients. Such probabilistic reference standards have been suggested or can be imagined in various settings. This paper aims at identifying conditions for unbiased estimation of sensitivity and specificity when a probabilistic reference standard is used in a diagnostic accuracy study. METHODS: The conditional distribution of the index test given the probabilistic reference standard carries the information on sensitivity and specificity. An explicit expression for this distribution is derived. It includes three different model components reflecting a potential association of sensitivity with the probabilistic reference standard, a potential association of specificity with the probabilistic reference standard, and a potential mean miscalibration of the probabilistic reference standard. Simple parametrizations of the three different model components are suggested. The dependence of the conditional distribution on the parameter values is investigated, and the bias of model-based and model-free estimates is investigated in a simulation study. RESULTS: Due to identifiability issues, it can only be expected to be able to include one of the three components in a model-based estimation. If absence of the other two components can be assumed, model-based estimation allows estimates with negligible bias. Model-free estimates show a high risk of bias. CONCLUSION: The use of probabilistic reference standards in diagnostic accuracy studies is challenging. Unbiased estimation of sensitivity and specificity can only be expected if two out of the following three conditions can be ruled out: 1) Association of sensitivity with the probabilistic reference standard. 2) Association of specificity with the probabilistic reference standard. 3) Insufficient mean calibration of the probabilistic reference standard. If this is the case, adequate statistical methods allow estimation with negligible bias. Arguing for absence of two conditions requires corresponding reasoning. PLAIN LANGUAGE SUMMARY: Evaluating the accuracy of a diagnostic test requires a systematic comparison of the test results with the true status in a series of patients. Determining the true status by a so-called reference test can be challenging in some patients. It has been suggested that the reference test should then result in a probability instead of a definite decision about the true status. It is unknown how to make use of such a probabilistic reference standard in analysing the diagnostic accuracy of the index test - i.e. the test of interest. This paper investigates how to estimate two key accuracy parameters - sensitivity and specificity - based on a probabilistic reference standard. Three conditions which challenge an unbiased estimation are identified: 1) Patients with a true positive status but difficult to diagnose for the index test are assigned lower probabilities than those easy to diagnose. 2) Patients with a true negative status but difficult to diagnose for the index test are assigned higher probabilities than those easy to diagnose. 3) The probabilistic reference standard systematically over- or underestimates the true status of the patients. If two of the three conditions can be excluded, it is possible to obtain unbiased estimates by adequate statistical methods. It is recommended to carefully discuss the potential presence of the three conditions whenever a probabilistic reference standard is used.
Werner Vach (Sat,) studied this question.