PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 1, 20260 citations

A pilot study on the impact of large language model assistance on the evaluation of complex medical living kidney donor candidates.

View Full Paper
HMH S MeshramCBC BhagatBMBhavin Modasia

Key Points

  • To evaluate the impact of large language model assistance on the assessment of complex living kidney donor candidates.
  • Prospective pre-post pilot study involving 14 physicians evaluating 30 complex kidney donor vignettes.
  • Primary outcomes measured included diagnostic accuracy, justification quality, and decision confidence.
  • LLM assistance provided via OpenAI ChatGPT-4o.
  • LLM assistance improved mean diagnostic accuracy by 0.27 and justification quality by 10.45.
  • Improvements in accuracy were recorded across both fellows (+0.29) and consultants (+0.25), with no role interaction.
  • Thematic richness increased notably, particularly in guideline-based and ethical framing areas.

Abstract

Background: Large language models (LLMs) are undergoing exploration as clinical decision support tools. However, their role in complex, high-stakes transplant nephrology decisions, including potential living kidney donor candidate (pLKDC) evaluation, remains unclear. Methods: In this prospective pre-post pilot study, 14 physicians (seven fellows and seven early-career nephrologists consultants) evaluated 30 standardized, deidentified complex pLKDC vignettes, first unaided and then with OpenAI ChatGPT-4o assistance. The primary outcomes were diagnostic accuracy (scale, 0-1), justification quality (0-100), and decision confidence (1-5). Secondary outcomes included thematic richness and confidence-accuracy calibration. Results: LLM assistance improved mean accuracy by 0.27±0.07 and justification quality by 10.45±1.96. Gains were consistent across fellows (+0.29 and +10.75, respectively) and consultants (+0.25 and +10.14), without significant role interaction. The largest improvements concerned vignettes involving hypertension with albuminuria (+0.43 accuracy) and metabolic or infectious risk stratification. Protocol-driven scenarios displayed the greatest benefit; ambiguous ethical or psychosocial cases had smaller gains. Thematic richness increased, especially in guideline-based and ethical framing domains. Calibration improved for consultants (slope, -0.04 to 0.10) but declined for fellows (-0.17 to -0.40), indicating residual overconfidence among less experienced clinicians. Conclusions: In this pilot study of artificial intelligence support for pLKDC assessment, LLM assistance was associated with consistently improved accuracy, reasoning quality, and thematic completeness. Protocol-driven decisions displayed the greatest gains. These findings warrant further exploration of LLMs as research tools in transplant nephrology, while underscoring the need for human oversight in nuanced ethical judgments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Meshram et al. (2026) studied this question.

synapsesocial.com/papers/69f442d4967e944ac5566383https://doi.org/10.4285/ctr.25.0085
Ask AI
Helpful
Bookmark
Share
View Full Paper