A guidelines-integrated large language model achieved 90.05% accuracy in identifying high-risk severe aortic stenosis patients, surpassing a EuroSCORE II-based approach (50.23%; P<0.0001).
Observational (n=231)
Does a guidelines-integrated large language model improve procedural risk stratification in patients with severe aortic stenosis compared to a EuroSCORE II-based approach?
A guidelines-integrated large language model significantly outperformed a traditional EuroSCORE II-based approach in stratifying procedural risk for patients with severe aortic stenosis.
Mean Difference: -39.82 (95% CI -47.96–-31.68)
Absolute Event Rate: 90.05% vs 50.23%
p-value: p=<0.0001
Abstract Background Traditional risk calculators like EuroSCORE II may inadequately classify severe aortic stenosis (AS) patients by overlooking key comorbidities and anatomical variables, limiting their utility in guiding transcatheter aortic valve implantation (TAVI) versus surgical aortic valve replacement (SAVR). We developed a guidelines-integrated large language model (LLM) based on the 2021 ESC guidelines for valvular heart disease to improve risk stratification compared to EuroSCORE II. Methods We retrospectively studied 231 AS patients evaluated by a Heart Team for procedural risk (low vs. high) from January 2022 to December 2024. Clinical vignettes were created to simulate Heart Team discussions. The guidelines-integrated LLM (GPT‑4o, version 2024-08-06) classified risk using a Forest-of-Thought prompting method. Its performance was compared to a EuroSCORE II-based approach (low risk: EuroSCORE II 4% and age 75; high risk: EuroSCORE II 8%), assessing accuracy, sensitivity, specificity, and ROC area. Logistic regression was used to assess the relative importance of EuroSCORE II versus other clinical variables. A subanalysis evaluated the guidelines-integrated LLM with versus without explicit EuroSCORE II input. Results In identifying high-risk patients, the guidelines-integrated LLM achieved 90.05% accuracy (95% CI, 86.07 to 94.02), notably surpassing the EuroSCORE II-based method at 50.23% (95% CI, 43.58 to 56.87) (mean difference −39.82%; 95% confidence interval CI, −47.96% to −31.68%; p0.0001). For low-risk stratification, it again outperformed the EuroSCORE II-based model (90.05% vs. 85.97%; mean difference −4.07%; 95% CI, −7.93% to −0.21%; p=0.039). Comparing LLM variants with and without EuroSCORE II information showed a 7.69% mean accuracy gain (95% CI, 2.82% to 12.56%; p=0.002) when EuroSCORE II was omitted. Sensitivity, specificity, and ROC analyses were consistent with these findings (Figure 1). Logistic regression indicated that excluding EuroSCORE II did not significantly alter the LLM’s overall weighting of EuroSCORE II variables (Mann–Whitney p=0.34). However, the lower performance with EuroSCORE II appeared linked to overemphasis on a limited subset of predictors, notably pulmonary artery systolic pressure (odds ratio OR 1.70; p=0.007), age (OR 1.39; p0.001) and kidney disease (OR 7.64; p=0.032). In contrast, the guidelines-integrated LLM without EuroSCORE II maintained a balanced weighting across multiple variables, except for age (OR 1.62; p0.0001) and male gender (OR 1.11; p=0.038). Conclusion A guidelines-integrated LLM strategy leveraging ESC guidelines provided superior high and low-procedural risk stratification of patients with severe aortic stenosis compared to a EuroSCORE II-based approach. By encompassing a wider range of clinically relevant factors this approach may enhance both clinical decision-making and individualized patient management, potentially better identifying candidate for TAVI.
Garin et al. (Sat,) conducted a observational in severe aortic stenosis (n=231). Guidelines-integrated large language model (GPT-4o) vs. EuroSCORE II-based approach was evaluated on Accuracy in identifying high-risk patients (MD -39.82%, 95% CI -47.96 to -31.68, p=<0.0001). A guidelines-integrated large language model achieved 90.05% accuracy in identifying high-risk severe aortic stenosis patients, surpassing a EuroSCORE II-based approach (50.23%; P<0.0001).