• AI models using electronic health records and claims data show strong potential for early identification of individuals at high risk of lung cancer. • Across 15 studies, model discrimination varied widely (AUROC range, 0.66 – 0.96), with most demonstrating good to excellent performance. • Despite promising results, most studies exhibited high risk of bias in the analytical domain, particularly related to missing data handling and lack of calibration assessment. • External validation and reproducibility were rarely performed, limiting generalizability and clinical applicability of current models. • Future work should prioritize rigorous validation, transparent reporting, and clinically interpretable models to enable safe integration into lung cancer screening and prevention. The growing global burden of lung cancer, coupled with widespread adoption of electronic health records (EHRs), has accelerated interest in artificial intelligence (AI)–based approaches for early risk identification. By leveraging routinely collected clinical data, these models offer the potential to identify individuals at elevated risk before the onset of symptoms, creating new opportunities for targeted prevention and timely intervention. However, the robustness of these models and their readiness for clinical use remain unclear. We conducted a systematic review to evaluate the development, performance, and reporting quality of AI-based lung cancer prediction models derived from EHR or administrative claims data. PubMed, Embase, Scopus, and Web of Science were searched for English-language studies published between 2010 and 2025. Eligible studies applied AI methods to predict or detect lung cancer using exclusively electronic clinical data sources. Study selection and data extraction were performed independently by two reviewers. Methodological quality and risk of bias were assessed using the CHARMS and PROBAST frameworks. Fifteen retrospective studies met the inclusion criteria. Model discrimination varied substantially, with reported AUROCs ranging from 0.66 to 0.96. While most studies demonstrated low risk of bias in participant selection and outcome ascertainment, the majority exhibited high risk of bias in the analytical domain. Common limitations included inadequate handling of missing data, limited assessment of calibration, and absence of external validation. Although internal validation was frequently reported, transparency in model reporting and reproducibility was often insufficient. AI models using EHR and claims data demonstrate considerable promise for early lung cancer risk prediction. However, important methodological and reporting gaps currently limit their generalizability and clinical applicability. Future research should emphasize rigorous external validation, comprehensive calibration assessment, and clinically interpretable model design to support safe and effective integration into lung cancer screening and prevention strategies.
Islam et al. (Wed,) studied this question.