Background: Large language models (LLMs) are undergoing exploration as clinical decision support tools. However, their role in complex, high-stakes transplant nephrology decisions, including potential living kidney donor candidate (pLKDC) evaluation, remains unclear. Methods: In this prospective pre-post pilot study, 14 physicians (seven fellows and seven early-career nephrologists consultants) evaluated 30 standardized, deidentified complex pLKDC vignettes, first unaided and then with OpenAI ChatGPT-4o assistance. The primary outcomes were diagnostic accuracy (scale, 0-1), justification quality (0-100), and decision confidence (1-5). Secondary outcomes included thematic richness and confidence-accuracy calibration. Results: LLM assistance improved mean accuracy by 0.27±0.07 and justification quality by 10.45±1.96. Gains were consistent across fellows (+0.29 and +10.75, respectively) and consultants (+0.25 and +10.14), without significant role interaction. The largest improvements concerned vignettes involving hypertension with albuminuria (+0.43 accuracy) and metabolic or infectious risk stratification. Protocol-driven scenarios displayed the greatest benefit; ambiguous ethical or psychosocial cases had smaller gains. Thematic richness increased, especially in guideline-based and ethical framing domains. Calibration improved for consultants (slope, -0.04 to 0.10) but declined for fellows (-0.17 to -0.40), indicating residual overconfidence among less experienced clinicians. Conclusions: In this pilot study of artificial intelligence support for pLKDC assessment, LLM assistance was associated with consistently improved accuracy, reasoning quality, and thematic completeness. Protocol-driven decisions displayed the greatest gains. These findings warrant further exploration of LLMs as research tools in transplant nephrology, while underscoring the need for human oversight in nuanced ethical judgments.
Meshram et al. (2026) studied this question.