This editorial emphasizes that new cardiovascular risk markers must demonstrate improved discrimination and calibration, not just statistical significance, to provide true prognostic value.
Risk assessment is a fundamental component of primary and secondary prevention of cardiovascular disease. A small set of cardiovascular risk factors, which we term the standard risk factors, consisting of gender, age, blood pressure, total (or low-density lipoprotein) cholesterol, high-density lipoprotein cholesterol, smoking behaviour and diabetes status, consistently show themselves to be important and independent predictors of cardiovascular disease [1–7]. The Framingham Study investigators, and other investigators, moving beyond considering one risk factor at a time, have incorporated this set of standard risk factors into multivariate risk functions that can estimate the absolute cardiovascular disease risk [1–8], simultaneously taking all these factors into account. A usual format is to estimate, for a given person, the absolute risk as a 10-year risk or probability of cardiovascular disease; for example, the risk of coronary heart disease (CHD) [1–5,7], stroke [6] or the combined CHD or stroke risk [2,8]. The search for other risk factors or markers that can add independent prognostic value beyond the above multivariate risk assessment has become intense, with the metabolic syndrome complex [9], non-invasive and sub-clinical indicators of disease [10,11] and inflammatory markers such as C-reactive protein (CRP) [12,13] leading the way. Statistical significance of the new variables in risk equations is usually attained, but prognostic improvement, as measured by increased discrimination between cardiovascular cases and non-cardiovascular disease cases and improved calibration (observed event rates matching predicted rates), is not often achieved [13]. Methods to evaluate discrimination and calibration in the setting of risk prediction exist [14,15]. The problem is that the new proposed risk factors usually do not show any improvement when subjected to these criteria. In this issue of the journal, Olsen et al. [16] take up the challenge of seeking variables that might improve prognosis beyond the small set of standard risk factors mentioned above. In hypertensives, they find that baseline and in-treatment reduction of the urine albumin/creatinine ratio (UACR) and electrocardiographic left ventricular hypertrophy (LVH) do attain statistical significance in multivariate models beyond gender, age, a Framingham risk score that contains the standard risk factors [3], smoking, existing cardiovascular disease and blood pressure and cholesterol levels. They make no attempt to evaluate whether discrimination is improved over the Framingham risk score, nor do they discuss the calibration of their model. Their study gives us the opportunity and setting to discuss the issues with respect to evaluating the prognostic improvement associated with a new proposed variable or risk factor. Four initial procedural decisions must be made in the quest for the evaluation of a new risk factor. They comprise: (i) the population of interest; (ii) the outcome of interest; (iii) how to incorporate the competing pre-existing set of risk factors; and (iv) selecting the appropriate statistical model and tests. The interplay of the four decisions impacts greatly on the ability to interpret meaningfully the results of the analysis. The population of Olsen et al. [16] comprises the set of hypertensives from the LIFE study. The major outcome is a composite of cardiovascular events [fatal and non-fatal myocardial infarction (MI) and stroke]. The pre-existing risk factors enter as a mixture of a Framingham risk score and individual risk factors, such as age, gender, blood pressure, cholesterol and smoking. Furthermore, other risk factors, such as existing cardiovascular disease, are included. The statistical model is a Cox proportional hazard regression with 3–4 years of follow-up after 1 year of anti-hypertensive treatment with either an atenolol- or losartan-based regimen. The investigators' initial procedural decisions potentially create analysis and interpretation problems. These problems relate to the use of the Framingham risk score as a means of controlling for the standard cardiovascular risk factors. The selected Framingham score [3] is a primary events model developed to estimate risk of an initial CHD event in individuals free of CHD. Here, CHD includes coronary death, MI, coronary insufficiency and stable angina. Using this Framingham risk score in a statistical analysis containing subjects with existing cardiovascular disease and predicting an outcome event that is a narrower set of coronary events (i.e. MIs) and also strokes (the latter an outcome not considered in the selected Framingham risk score) is inconsistent with the score and may only produce incomplete adjustment of the standard risk factors. A more useful analysis probably would have comprised using a Framingham risk score dealing closer with the outcome event that was the outcome under consideration in the analysis. These do exist (e.g. for primary cardiovascular events [2]; a combination of [5] and [6] for primary CHD and stroke, respectively; and for secondary events [4]). The lack of use of an appropriate risk score for the population and outcome under consideration leaves open the question of whether the correct adjustments have been made. At least, the appropriate interpretation of what is the true effect of the standard risk factors in the final model is unclear. The Framingham risk score is a composite for capturing risk. The use of a composite risk score such as the Framingham score can be very useful in the context of multivariate modeling. If performed correctly, the coefficient of the score in the final multivariate model provides a summary index of the importance of the composite impact of the entire set of variables contained within it. It quantifies directly the composite hazard ratio or relative risk. However, when used as a predictor or control variable in the model, the selected Framingham score should be, at a minimum, consistent with the population under consideration and the outcome of interest. Furthermore, secondary analyses employing the individual variables of the composite score should also be performed. For the study of Olsen et al. [16], a secondary analysis directly considering the risk factors gender, age, blood pressure (initial and 1 year change), cholesterol (initial and 1 year change), smoking, diabetes and existing cardiovascular disease in the analysis would have been very useful for understanding the independent effect of UACR and LVH. In this analysis, the effect of each risk factor could be evaluated directly without the possibility of it becoming confused in the risk score. Given that the above four initial decisions have been made correctly, the next step is to interpret correctly the prognostic value of the new variable in the multivariate model that adjusts for the standard risk factors and other relevant variables, such as treatment and existing disease status. Two preliminary requirements must be met. Evaluation of the statistical significance of the variable and the quantification of relative risk (or hazard ratio) are essential. As a prerequisite to declaring the independent prognostic improvement (or the ability to add to risk prediction), the new variable must be statistically significant and its relative risk must be clinically meaningful. For example, the latter should correspond to at least a two-fold increase when comparing the first to the last quartile. Depending upon the situation, clinical meaningfulness will vary. The important point is that it needs to be considered. Mere statistical significance is necessary but not sufficient. The same is true for clinically meaningful hazard ratios. Many new potential markers and variables will pass the statistical significance requirement. Some will have clinically meaningful hazard ratios. The next hurdle is whether the new variable has any additional prognostic value in risk prediction. This moves us into the discussion of added discrimination and good calibration. Discrimination relates to the ability to rank people correctly in order of their risk and to separate those who will develop an event from those who will not. It is usually measured by the C statistic of the receiver operating characteristic curve [14,15]. Calibration relates to the ability of the model predictions to match the actual observed rates. Calibration is measured by a chi-squared test which compares observed event rates to model predicted rates [14]. Recently, these model performance measures have been developed for survival models such as the Cox regression [5,14,15]. For a new variable to make an independent prognostic improvement (or equivalently improve risk prediction), its addition to a risk score function should significantly improve discrimination and maintain good calibration. When these criteria are applied to the evaluation of the prognostic usefulness of a variable, many new constructs and markers, such as the metabolic syndrome, and many of the new markers, such as CRP, do not indicate any usefulness beyond the standard risk factors. Olsen et al. [16] do not present any analysis relevant to discrimination and calibration. The translation of the statistical significance they have found into improved prognosis is unanswered. How can a new variable add to discrimination? Unfortunately, it appears that, in terms of hazard ratios, a new marker needs to be extremely important to improve discrimination [17] by an order of magnitude exceeding 2 or 3. New work in clinical thinking and mathematical statistics is needed to indicate the nature of variables that can improve discrimination. The relationship of the new marker with standard risk factors is needed. It has to offer some correlation with the outcome that is beyond the standard risk factors. It cannot be highly correlated with the set of standard risk factors. The identification of subgroups (such as those at some specified risk level) where the new variable has special and meaningful added importance is also needed. That Olsen et al. [16] focus on hypertensives is a promising avenue. The statistical methods to evaluate and quantify this added prognostic gain still remain to be developed. Is improved prognosis related to new risk factors or the quantification of disease states or treatment effects? A major motivation for developing the Framingham risk functions was to assess the importance of modifiable risk factors such as blood pressure, cholesterol and smoking in the presence of the background conditions of gender and age and the ‘disease state’ of diabetes and presence of LVH [1,2,4,6]. Inflammatory markers such as CRP fit into the framework of a potentially modifiable risk factor. The results of non-invasive tests such as carotid ultra-sound are measures of disease states. If it adds independent prognostic value then it is because it has quantified an existing disease state. The reduction in albuminuria in the study by Olsen et al. [16] is different from both of the above. This reduction reflects a treatment effect. To understand the prognostic usefulness of this new proposed risk variable, it is necessary to evaluate first the treatment effect on the basic risk factor, namely blood pressure. Blood pressure control may comprise the important variable that results in a reduction of cardiovascular events. The reduction in UACR may be a consequence of blood pressure control and a mechanism or pathway for a reduction in cardiovascular events. Olsen et al. [16] did not find blood pressure to be significant. The use of the Framingham function in the analysis may have confounded their ability to understand the entire impact of blood pressure reduction. The inclusion of blood pressure and its control (change in blood pressure) directly in the analysis might have produced a clearer picture. In conclusion, Olsen et al. [16] have attempted to address the important question of how treatment effects can improve prognosis. The authors performed solid analyses to control for the standard risk factors and other time-varying conditions. We congratulate them on their efforts and identify issues and problems with their analyses. We recommend further analyses that address the problem of Olsen et al. [16] plus the more general issue of establishing and quantifying the independent prognostic value of a new variable in the risk scoring setting. We have also attempted to identify deficiencies with respect to present clinical and statistical thinking and methods. We strongly urge further developments that can formalize the methods and thinking behind this evaluation.
No takes yet. Share an insight, caveat, or question.
Ralph B. D’Agostino (2006) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: