Abstract Rationale Venous thromboembolism (VTE) is the third most common thrombotic vascular disease worldwide. The Caprini and Padua scores are commonly used risk assessment tools for VTE, however, current way of manual assessment faces limitations of efficiency, accuracy, and cross-center consistency. We aimed to develop an automated VTE risk assessment framework (Padua and Caprini) for VTE based on base large language model (LLM) with expert-augmented prompt and verify its feasibility and stability in practice. Methods We utilized data from a nationwide, multicenter, prospective cohort, including patients≥18 years with anonymized electronic medical records (EHR) from 30 hospitals across China. The system prompts for the extraction of Padua and Caprini scores were developed based on original literature and judgement from 19 attendings with over 20-year experience. Fifty cases from 10 hospitals and 200 cases from another 20 hospitals were selected as test set (for prompts optimization, preliminary accuracy estimation and LLM evaluation) and external validation set (for comprehensive evaluation), respectively, through stratified random sampling method. During LLM extraction, the chain of thought (COT) was developed, and rationales were required. To evaluate the performance of LLM in terms of accuracy, kappa coefficient and efficiency, the gold standard was established by two attendings with over 10-year experience. Results The overall extraction accuracy for Padua score was 0.99 and Kappa was 0.94; for Caprini score, the overall accuracy was 0.98 and Kappa was 0.92. Both scores showed high consistency with the gold standard. LLM performed excellently in identifying most risk factors in the Padua score (all accuracy and Kappa coefficients0.90) except for “limited activity”. For the Caprini score, most variables showed good agreement (Kappa coefficient≥0.9), beside items requiring further temporal or contextual reasoning (e.g., “major surgery”, “bed rest”). In risk stratification, the LLM-based automated scores were highly consistent with the gold standard: for the Padua score, the accuracy (0.96) and Kappa (0.90); for the Caprini score, LLM achieved high accuracy (0.86) and moderate consistency (Kappa 0.64). Additionally, LLM significantly improved efficiency: the average automated scoring time per case was about 10-20s for Padua and 58-80s for Caprini, much faster than manual evaluation. Conclusion Through expert-knowledge augmented prompt engineering, base LLM achieved expert-level performance on the Padua score without any fine-tuning and high accuracy and consistency on the Caprini score. This approach significantly improved scoring efficiency, enabling automated VTE risk assessment from EHRs and establishing a universal, replicable, and resource-friendly automated scoring paradigm. This abstract is funded by: None
Fan et al. (Fri,) studied this question.