Methodological review demonstrates discrimination assessment using the C-index in prognostic risk models, highlighting why strong risk associations do not ensure added predictive accuracy.
The objective of all diagnostic and therapeutic decisions taken by physicians in the care of patients is that of improving prognosis by preventing or modifying the natural evolution of the disease. Therefore, prognostic research is an important area of investigation in clinical epidemiology. A clinician usually formulates a prognosis on the basis of patient characteristics or biomarkers, i.e. clinical indicators of various sort (for example serum creatinine, arterial pressure, left ventricular mass as measured by echocardiography, etc.), reflecting normal or pathological pathways related to the exposure–outcome axis or to biological responses to a therapeutic intervention. Although biomarkers can be classified into different categories, here we only consider biomarkers of exposure. An ideal prognostic biomarker should allow the early identification of individuals at risk for a given outcome, and should be relatively easy to measure with acceptable costs. Before being introduced into clinical practice, biomarkers need to be properly validated. Formal proof that the use/modification of the same biomarker is able to affect/predict the prognosis at an individual or population level is a fundamental step in the validation process of a biomarker. Multivariate modelling of pertinent clinical outcomes by the combination of multiple biomarkers and other indicators allows the generation of risk prediction scores. The most popular of these risk scores is the Framingham risk score (FRS), a score based on demographic information (age, sex), an environmental exposure (smoking) and three biomarkers (systolic pressure, total cholesterol and high-density lipoprotein (HDL) cholesterol). FRS estimates the individual risk of myocardial infarction and coronary death in 10 years. Accuracy and generalizability are important issues related to the use of prognostic biomarkers and risk prediction scores as well. Accuracy is the agreement between the outcome predicted by the biomarker and the actual occurrence of the outcome. Generalizability is the capacity of the same biomarker to provide accurate predictions in population samples different from that in which the biomarker was originally validated. There are three commonly used methods to assess the accuracy of biomarkers for predicting clinical outcomes: discrimination, calibration and reclassification. In this article, we focus on discrimination, and in the next one, on calibration and reclassification. The C-index may take values ranging from 0.5 (no discrimination) to 1.0 (perfect discrimination). In our instance, the C-index predicted on the basis of risk tables is 0.62 which implies that, in a hypothetical experiment in which we randomly select pairs of individuals with and without CV events, the 5-year probability of CV outcomes estimated by risk tables will be higher (62% of times) in individuals with than in those without CV outcomes. More precisely, the C-index is the proportion of pairs of subjects (with opposite outcome), where the one who actually experiences the adverse outcome had a higher (predicted) probability of event. To be statistically significant, the C-index should have a 95% confidence interval not including 0.5. In our example, also due to the small sample size, a C-statistics of 0.62 is not significant because the corresponding 95% confidence interval (CI) includes 0.50 (95% CI ranging from 0.19 to 1.00). The C-statistics is conceptually similar to receiver operating characteristic (ROC) curve analysis [3]. Example of C-index calculation (a measure of discrimination). See text for details. In a recent paper, Wang and co-workers [4] investigated the incremental value of the simultaneous use of multiple biomarkers for predicting the probability of death in a cohort of 3209 individuals attending a routine examination cycle of the Framingham Heart Study. Although in this paper the authors used a slightly modified C-statistics [5], the interpretation of the C-index is identical to that of Example 1. The authors identified five biomarkers significantly related to the mortality risk: brain natriuretic peptide, C-reactive protein, urinary albumin-to-creatinine ratio, homocysteine and renin. By combining these biomarkers, the authors calculated an individual multimarker score (for details see Reference [4]). Then, they stratified the study population into three risk groups according to the value of this biomarkers score: low, intermediate and high risk. In a Cox's model including traditional [age, sex, arterial pressure, total and low-density lipoprotein (LDL) cholesterol, smoking, diabetes and cardiovascular co-morbidities] and non-traditional risk factors (body mass index and serum creatinine) as well as the multimarker score, individuals in the high risk category (those with a high multimarker score) had a relative risk of death that was four times higher than that of individuals in the low risk category (relative risk: 4.08, 95% CI: 2.51–6.62, P < 0.001) indicating that the multimarker score provided risk stratification above and beyond that provided by standard or conventional risk factors. The incremental value of the multimarker score for discriminating individuals who died from those who survived was also investigated by C-statistics. This analysis showed that a prediction model including age, gender and other conventional risk factors but not the multimarker score gave a C-index for death of 0.80, a figure that did not differ from that achieved by a model including age, sex and the multimarker score (C-index = 0.79). The authors concluded that, although the multimarker score was able to stratify the risk of death in the study cohort [relative risk (high versus low multimarker score): 4.08, P < 0.001], it had no additional predictive value for discriminating individuals who died from those who survived. This apparent discrepancy is due to the fact that relative risk and C-index provide quite different information. The relative risk is a measure of effect particularly useful in aetiological research (that is, it serves to assess/explain the strength of a specific biomarkers–outcome relationship) while the C-index is a measure of accuracy well suited for pure prognostic research (that is, it serves to ascertain the predictive ability of a given biomarker without any concern about aetiology) [6]. In this regard, it is important to note that the discriminatory ability of one or more biomarkers for a given outcome strictly depends on the clinical/epidemiological context being considered as well as on the structure of the prediction model being tested. The C-index is potentially insensitive in assessing the impact of adding a new predictor to a previous prognostic model when other strong predictors of the outcome are already included into the same model [7]. Discrimination reflects the ability of a prognostic model to correctly identify a clinical status (that is a condition having only two possibilities: event/non-event; died/survived, etc.). Discrimination is of particular importance for clinicians because they are generally interested in knowing how much a biomarker or a prediction score is able to distinguish between individuals at high risk for a given event from those not being at high risk (the discriminating power of the prediction). Transparency declaration. The results presented in this paper have not been published previously in whole or part, except in abstract format. Conflict of interest statement. None declared.
No takes yet. Share an insight, caveat, or question.
Tripepi et al. (2010) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: