Abstract Background Accurate identification of high-risk individuals is crucial for personalised prevention services. Cardiovascular diseases (CVD) are a major health burden globally, and cholesterol-lowering medications are one of the main interventions for prevention. Traditional risk models require clinical visits and blood tests, but little research has been conducted on the feasibility of utilising existing electronic health records (EHR) to predict CVD risk. Estonian EHR offers a unique opportunity to develop and validate predictive models without requiring additional patient visits. A reliable risk prediction tool could enable more targeted and cost-effective CVD screening programs. Purpose This study aimed to develop and internally validate two CVD risk prediction models for estimating the 5-year risk of primary cardiovascular disease in men and women. The models are based on available EHR data, demonstrating the feasibility of using existing health records to improve risk assessment. Methods We conducted a retrospective cohort study on a 10% random sample of the Estonian population (n = 150,824) from 2012 to 2019, incorporating data from three national health databases: claims, prescriptions, and EHRs. The analysis included 52,172 individuals aged 40-74 years, all free of cardiovascular disease at baseline. We created two models using the Cox proportional hazards model: one with age and gender, and a more complex model incorporating age, gender, and a range of clinical factors (e.g., diabetes, hypertension, atrial fibrillation). We considered continuous variables such as systolic blood pressure, cholesterol, and body mass index, but missing data for these variables exceeded 50%. Both models were validated using Harrell’s C statistic and 10-fold cross-validation. Net reclassification improvement was calculated to compare the performance of the two models. Results Among the 52,172 participants, we identified 1935 incident cases of CVD. The age-gender model had a C statistic of 0.750 (95% CI 0.738–0.761) in the training set and 0.747 (95% CI 0.727–0.768) in the test set. The more complex model showed a C statistic of 0.762 (95% CI 0.751–0.773) in the training set and 0.757 (95% CI 0.737–0.778) in the test set, demonstrating improved predictive performance. The net reclassification improvement was +0.112 (95% CI 0.098–0.123), indicating significant enhancement in risk classification when including additional clinical factors. Conclusions We developed two robust CVD risk prediction models, which demonstrate the potential of using only EHR data for accurate risk assessment. The inclusion of additional clinical factors improves predictive performance, allowing healthcare providers to identify high-risk individuals more efficiently. This work lays the foundation for more personalised and cost-effective CVD screening strategies, with potential applications in global public health initiatives.Baseline characteristics Reclassification analysis
Loo et al. (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: