PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Annals of Plastic Surgery0 citations

Evaluating GPT-4 Performance on Plastic Surgery Oral Examination Vignettes

View Full Paper
JSJason SalvatoJSJordan SalvatoYBYasmeen M. Byrnes

Key Points

  • The study aims to assess GPT-4's ability to address clinical vignettes relevant to plastic surgery and its potential in educational contexts.
  • Input of twelve clinical vignettes into GPT-4 covering various plastic surgery domains.
  • Scoring of responses by two board-certified plastic surgeons based on six clinical domains.
  • Assessment of interrater reliability using weighted κ.
  • Conducted a structured qualitative interview to gather additional insights.
  • GPT-4 achieved mean scores from 2.3 to 3.7 across the clinical domains evaluated.
  • Passing rates were 91.7% and 83.3% for the two raters, respectively.
  • Diagnosis and Complication Handling were the highest-performing domains, while Case Introductions had the lowest scores.
  • Interrater agreement was strong, with κ values between 0.70 and 0.89.
  • Qualitative findings highlighted GPT-4's strengths and weaknesses in various clinical scenarios.

Abstract

Background The use of artificial intelligence (AI) in medicine is rapidly evolving. However, its role in plastic and reconstructive surgery remains underexplored. Plastic surgery requires nuanced, dynamic decision making, and individualized care, making AI integration challenging. This study evaluates GPT-4's ability to respond to American Board of Plastic Surgery (ABPS) Oral Board-style clinical vignettes and assesses its potential as a decision making and educational adjunct. Methods Twelve clinical vignettes spanning aesthetic, reconstructive, hand, craniofacial, pediatric, and trauma surgery were input into GPT-4. Each response was scored by 2 board-certified plastic surgeons across 6 clinical domains: Case Introduction, Diagnosis, Treatment Planning, Patient Counseling, Operative Steps, and Complication Management. Domains were scored on a 0–4 scale; scores <2 were considered failing. Interrater reliability was assessed via weighted κ . A structured qualitative interview was conducted. Results GPT-4 achieved mean domain scores ranging from 2.3 to 3.7 across vignettes, with passing rates of 91.7% (rater 1) and 83.3% (rater 2). The highest-performing domains were Diagnosis (84.4%) and Complication Handling (82.3%), and the lowest-performing domains were seen in Case Introductions (62.5%). Interrater agreement was strong across domains ( κ = 0.70–0.89). Qualitative findings emphasized GPT-4's accuracy in guideline-based cases, limitations in trauma triage, and potential as an educational resource. Conclusion GPT-4 demonstrates concordance with board-level clinical reasoning in structured plastic surgery scenarios. However, limitations in heuristic judgment and real-time adaptability underscore the need for optimization with specialty-specific training datasets and performance rubrics. With appropriate curation, GPT-4 could serve as a valuable supplement in surgical education and oral board preparation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Salvato et al. (2026) studied this question.

synapsesocial.com/papers/698d6e5a5be6419ac0d54028https://doi.org/10.1097/sap.0000000000004671
Ask AI
Helpful
Bookmark
Share
View Full Paper