Key points are not available for this paper at this time.
Design
Editorial
Researchers should carefully consider the clinical implications of false-positive and false-negative errors before recommending a biomarker cut-point, rather than relying solely on mathematical maximization like Youden's index.
When you come to a fork in the road, take it. —Yogi Berra Assessing the ability of a biomarker or predictor variable to decide whether or not patients have a certain condition, or to predict those who are likely to contract it, is a vital component of perioperative research. In such studies, researchers commonly conduct a receiver operating characteristic (ROC) analysis in which they plot the sensitivity versus 1 minus the specificity of the outcome for each observed value of the predictor variable, where sensitivity is 1 minus the false-negative probability and specificity is 1 minus the false-positive probability. One then reports the “area under the ROC curve” (AUC) of that plot as a measure of how well the predictor variable can discriminate those who have the event of interest from those who do not. In such an analysis, an AUC of 0.50 would be attained by randomly guessing at whether a patient had the outcome or not, and 1.0 would indicate a perfect discrimination between events and nonevents by using the predictor variable. If there is good discriminative ability, such that the estimated AUC is noticeably >0.50, one might then search for a cut-point of the predictor either: (1) to be used in decision-making on whether a patient has or is likely to contract the condition; and/or (2) to determine whether a more invasive test should be conducted or not. So how do we best identify the optimal cut-point for such a predictor? In this issue of Anesthesia & Analgesia, Gomez Builes et al1 sought to find a cut-point in trauma patients for the value of maximum lysis (ML), below which a patient would be less likely to survive 48 hours. They used logistic regression, with mortality as the outcome and ML as the predictor variable, and estimated the AUC for ML as a predictor of mortality (although the observed AUC is not reported). They then used Youden’s index, which maximizes the sum of sensitivity and specificity, to choose the best cut-point for ML in discriminating which patients would be expected to die within 48 hours. Youden’s index (J) is calculated as J = sensitivity + specificity − 1 for each observed value of the predictor. The value of the predictor with the largest J value is chosen as the best cut-point.2 Thus, Youden’s index has a maximum value of 1.0 and a minimum value of 0. As pointed out by Zhou et al,3 Youden’s index is also equal to sensitivity minus the false-positive rate, and reflects the likelihood of a positive result among patients with versus without the condition. Gomez Builes et al1 found that a value of ML <3.2% maximized Youden’s index, and they, therefore, chose ML <3.2% to indicate “fibrinolysis shutdown” for a patient. Using this cut-point, they observed a sensitivity of 42% (95% confidence interval [CI], 27%–57%) and a specificity of 76% (95% CI, 51%–88%) for predicting mortality within 48 hours. However, the striking gap between sensitivity and specificity in Gomez Builes et al1 should prompt one to ask whether Youden’s index is a good method to define a best cutoff—at least in this situation—and to question whether there are other available methods. For example, if sensitivity and specificity are equally important, would not the optimal cut-point have sensitivity and specificity that are more similar?3 And if they are not equally important, should that not be reflected in the choice of an optimal cut-point? Finally, are there times when no cut-point is worthy of being published and used in practice? And how can we tell? FORK IN THE ROAD There are, in fact, several methods that can be used to determine a best cutoff.4 Perhaps most commonly, the chosen best cut-point is the one that maximizes both sensitivity and specificity (not the sum, as in Youden’s index). This value is the point on the ROC curve at which sensitivity and specificity are equal, or as close to equal as possible for the given data. Maximizing both sensitivity and specificity assumes that sensitivity and specificity are equally important, and so this is a good method to use if that is the prerequisite, or if it is not known which is more important. Another option would be to choose the point that has the smallest distance from the upper left corner of the ROC curve. For a strong predictor, this value will often coincide with or be very close to the first method: that of equalizing sensitivity and specificity. Finally, Youden’s index, which maximizes the sum of sensitivity and specificity, might be used, as by Gomez Builes et al.1 Cut-points chosen using Youden’s index would also tend to be closer to the method of equalizing sensitivity and specificity when the predictor is very strong, corresponding to an AUC that is high (eg, >0.80). But that was apparently not the situation with Gomez Builes et al1 because the estimated sensitivity was <0.50, and an enormous 0.34 lower than specificity. Gomez Builes et al1 claim that because Youden’s index does not, by definition, favor either sensitivity or specificity, it might be useful when neither is known or assumed to be more important. This is spurious because when a predictor is not strongly associated with the outcome, such that the AUC is moderate at best (eg, AUC <0.80), Youden’s index can produce “best” cut-points that have very different sensitivity and specificity, such as in their own data set. The best approach in such a situation might be to not choose any best cut-point. However, if one is chosen, it would optimally be one with similar sensitivity and specificity. Gomez Builes et al1 correctly point out that in situations such as theirs, in which a new clinical decision-making tool is being introduced into practice to predict a severe condition, a higher specificity is generally preferred over high sensitivity because one is interested in reducing false-positive findings. It appears that Gomez Builes et al1 were fortunate in that regard because Youden’s index could have just as easily given a cutoff with considerably higher sensitivity than specificity. Therefore, in situations in which either sensitivity or specificity is more important, or when they are deemed equally important, Youden’s index is not the best choice and should probably not be used. Actually, it is difficult to see when Youden’s index would be the best option, given that other methods achieve the goal of providing the best available cut-point without the limitations. STUDY DESIGN When seeking an optimal cut-point, researchers should specify in the design phase that one will only be chosen and recommended for use in practice if it meets certain criteria. A reasonable first step would be to state the minimally acceptable value of AUC, such as 0.75, before searching for a best cut-point for the predictor. Note that an AUC of 0.75 is 50% of the way between 0.50 (random guessing) and 1.0 (perfect discrimination). However, it still might be possible to find a cut-point that meets the study criterion for sensitivity and specificity with a lower AUC. Next, authors should a priori define minimal criteria for sensitivity and specificity, like each being at least 70% or 75%, before a cut-point would be recommended in practice. Alternatively, authors might specify, for example, that sensitivity needs to be at least 90% and specificity at least 65%, or vice versa, depending on whether false-negative or false-positive errors are more problematic for the clinical application.3 If, in a given study, making either false-positive or false-negative decisions are clinically worse, the optimal cut-point would not be one that maximized both sensitivity and specificity. Given the particular specifications, the ROC curve can be searched to maximize a potential cutoff while fulfilling the clinical requirements. If, indeed, sensitivity and specificity are deemed equally important, the method of searching the ROC curve to maximize each of them seems extremely useful, and it would avoid the pitfalls of Youden’s index. In the Figure, there are 2 exemplary ROC curves representing biomarkers A and B, each predicting a binary disease condition. Notice that for biomarker A, Youden’s index (J) is highest for the biomarker value that has a sensitivity of 0.55 and a specificity of 0.85, for J = 0.55 + 0.85 − 1 = 0.40. But Youden’s index is 0.30 for a biomarker value that has equal sensitivity and specificity of 0.65.Figure.: ROC curves for biomarkers A and B, each predicting a binary disease condition. Data in parentheses by each point are sensitivity and specificity. Biomarker A has its highest Youden’s index J (sensitivity + specificity − 1) of 0.40 at a biomarker value with a sensitivity of 0.55 and a specificity of 0.85, and a Youden’s index is 0.30 for the value that equalizes sensitivity and specificity at 0.65. On the other hand, the optimal cut-point for biomarker B is clearly the point that equalizes sensitivity and specificity at 0.75 (Youden’s index of 0.50), which would often be considered sufficiently strong diagnostic accuracy for practical use. Simply using the largest observed Youden’s index to determine a best cut-point is not generally recommended. ROC indicates receiver operating characteristic.Depending on what values of sensitivity and specific were identified a priori as being minimally sufficient, it might be that neither cut-point has an adequate combination of sensitivity and specificity to be sufficiently useful in practice. On the other hand, maximizing both sensitivity and specificity for biomarker B yields estimates of 0.75 for each, which, for many applications, would be considered sufficiently strong diagnostic accuracy for use in practice. In some situations, investigators want to restrict their estimate of diagnostic accuracy to a certain value or range of specificities, and not consider the entire range of the predictor or the ROC curve. For example, they might want to estimate the “partial” AUC when specificity is limited to between 0.70 and 0.90, or they might want to only estimate the sensitivity at a particular value of specificity. In such scenarios, the search for a best cut-point would be limited to the desired area of inference.3 Not uncommonly, the estimated AUC is low, and it is difficult to obtain any cut-point with adequate sensitivity and specificity for clinical practice. In such situations, there will certainly be regions in which either sensitivity or specificity may be adequate, but not both—such as for biomarker A in the Figure. Researchers should not shy away from or avoid concluding that the relationship between the biomarker and outcome is not strong enough to recommend any particular cut-point for practice. To make a recommendation when diagnostic accuracy measured by sensitivity and specificity is weak may put patients at risk, lead to unnecessary testing, or at least to poor decision-making. Finally, whenever an optimal cut-point is estimated from a study, as with any other estimate, it should be accompanied by a CI. A CI for a best cut-point can be estimated using bootstrap resampling with replacement, which involves randomly resampling the data many times, and each time estimating a best cut-point, then obtaining either the standard error or the CI from the distribution of those observed cut-points.5 In conclusion, the search for an optimal cut-point to predict outcome in a diagnostic accuracy study should be done only after careful consideration of the relative costs or adverse ramifications of false-positive and false-negative errors. Researchers should recommend cut-points for use in practice only if diagnostic accuracy is sufficient for the given situation, rather than blindly maneuvering that proverbial fork in the road. DISCLOSURES Name: Edward J. Mascha, PhD. Contribution: This author helped design, write, and revise the manuscript; and analyze the data. This manuscript was handled by: Thomas R. Vetter, MD, MPH.
No takes yet. Share an insight, caveat, or question.
Edward J. Mascha (2018) studied this question.