PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 8, 2025International Journal of Surgery3 citationsOpen Access

Can large language models aid pre-surgical epileptogenic zone localization? A multi-source text analysis performance study in drug-resistant epilepsy

View Full Paper
YSYueqian SunSGShihao GeYWYangyang Wang

Key Points

  • Models achieved 100% accuracy in cohort two for localization of the epileptogenic zone, indicating strong predictive validity.
  • Performance in localizing the epileptogenic zone was benchmarked against surgically validated data from 154 patients with drug-resistant epilepsy.
  • Analysis included three leading LLMs evaluating medical records and clinical text to predict surgical outcomes effectively.
  • Larger scale validation may enable LLMs to improve clinical decision-support tools for epilepsy surgical candidates.

Abstract

Large language models (LLMs) show promise for biomedical text analysis, but their ability to localize the epileptogenic zone (EZ) from pre-surgical data is underexplored. We evaluated three leading LLMs on predicting the surgically validated EZ using multi-source unstructured clinical text from 154 patients (two cohorts) with drug-resistant epilepsy who achieved postoperative seizure freedom (Engel Class I). Models performed EZ laterality classification, probabilistic lobar localization, and SEEG stratification based on five presurgical text sources, with performance benchmarked against resected lobes. All models achieved 98.3% accuracy in Cohort 1 and 100% in Cohort 2 for laterality. In lobar localization, GPT-4.1 and Claude 3.7 Sonnet achieved a median score of 70, significantly higher than Deepseek-R1’s 60 (p 0.05). A modality ablation analysis revealed that GPT-4.1’s performance was most significantly impaired by removing medical records (p<0.001) and MRI reports (p = 0.014). Predictions exhibited excellent test-retest reliability (ICC = 0.951), and higher SEEG stratification scores were assigned to patients who underwent invasive monitoring (p < 0.001). In conclusion, LLMs can accurately infer the surgically validated EZ from raw preoperative text and may serve as valuable clinical decision-support tools.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sun et al. (2025) studied this question.

synapsesocial.com/papers/694020e22d562116f28fa9f8https://doi.org/10.1097/js9.0000000000004435
Ask AI
Helpful
Bookmark
Share
View Full Paper