Validation study reveals moderate broad-level but poor exact ICD code translation by GPT-4, suggesting current LLMs require refinement before deployment.
The International Classification of Diseases (ICD) is a medical coding system used for healthcare system administration, public health surveillance, and research. Crosswalk tables, which map diagnosis codes across ICD versions, are time-consuming and costly to produce but are essential to reduce medical coding errors when healthcare systems periodically adopt ICD updates. Automated crosswalk development could better support clinical staff during transition periods and facilitate research spanning multiple ICD versions. Our aim was to evaluate the ability of a pre-trained large language model (LLM) to generate ICD crosswalks. We evaluated the accuracy of the fourth-generation OpenAI Generative Pre-trained Transformer (GPT-4) model to translate chronic disease diagnoses across the 9th and 10th revisions of U.S. and Canadian ICD systems. Nine prompting strategies were developed. The three most-accurate prompts were combined to form composite prompts for Canadian and U.S. contexts. Accuracy was evaluated against crosswalks developed by Canadian and U.S. health services organizations. Each prompt was executed 10 times to assess variability, with mean accuracy ± standard deviation (SD) reported across replications. Across the nine prompting strategies evaluated for translating Canadian ICD codes, accuracy ranged from 32.5% to 47.4%, with the highest accuracy attained when prompts included diagnosis code labels. Prompting also impacted model variability, with SDs ranging from 0.6% to 3.6%. When combining the three best-performing prompts, GPT-4 achieved accuracies of 48.3% ± 0.8% (Canada) and 38.8% ± 0.6% (U.S.). Accuracy varied substantially across disease categories: 39.0%-70.0% (Canada), 30.1%-58.3% (U.S.). When evaluating only the first three characters of predicted codes, GPT-4 achieved 84.7% ± 0.6% (Canada) and 86.2% ± 0.4% (U.S.). LLMs are highly sensitive to prompting strategy, substantially impacting both accuracy and variability. Testing multiple prompting strategies is recommended when employing LLMs for clinical or research applications. GPT-4 achieved moderately high accuracy when predicting broad diagnostic categories (first three characters) but performed poorly when generating more specific diagnoses. While GPT-4 performance was insufficient for deployment, LLMs show promise for automated ICD code translation and could become reliable with targeted refinement.
No takes yet. Share an insight, caveat, or question.
Monchka et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: