PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 20260 citationsOpen Access

Human and artificial intelligence performance in radiographic caries detection: ex vivo tooth section-referenced evaluation and implications for clinical decision-making.

View Full Paper
CGC. GanßKJKatja JungLSLea Schilling

Key Points

  • This evaluation aims to compare the caries detection performances of an AI system and human raters using tooth section references.
  • Assessment of radiographs and tooth sections from extracted teeth with about 639 sites evaluated.
  • Three dentists and an AI system analyzed the images with graded outputs.
  • Performance was quantified using ROC, PR analyses, and calibration metrics.
  • Various decision approaches for early dentine lesions were compared using clinical-loss analysis.
  • The AI system showed higher sensitivity for approximal enamel lesions compared to raters (0.60 vs 0.42).
  • In approximal dentine lesions, the AI system also demonstrated higher sensitivity (0.84 vs 0.72) and high discrimination (AUC 0.90 vs 0.86).
  • For cervical dentine lesions, both AI and raters performed well, with the AI achieving sensitivity and specificity of 0.92 and 0.84.
  • Among decision strategies for D3 lesions, the AND approach yielded the lowest clinical loss.

Abstract

Objectives To evaluate the caries detection performance of a commercially available AI-system (Nostic) and human raters for radiographic caries detection using tooth sections as reference. Methods Radiographs and corresponding tooth sections from extracted teeth (548 approximal and 91 cervical sites) were assessed by three dentists and AI-system (graded/probabilistic outputs). Performance versus reference was quantified using ROC/PR analyses, calibration metrics, and threshold-based measures with bootstrapped confidence intervals. Rater-AI-system agreement/disagreement subsets were analysed. For early dentine lesions (D3), four decision approaches (rater-only, AI-system-only, rater-OR-AI-system, rater-AND-AI-system) were compared using a weighted clinical-loss analysis across varying false-positive penalties. Results For approximal enamel lesions (D1/2), AI-system was more sensitive than raters (0.60 vs 0.42) with slightly lower specificity (0.80 vs 0.84) and modest discrimination (AUC 0.70 vs 0.64). For approximal dentine lesions (D3/4), discrimination was high (AUC 0.90 AI-system vs 0.86 rater); The AI-system was more sensitive (0.84 vs 0.72) while raters were more specific (0.94 vs 0.88). For cervical dentine lesions (D3/4), both performed well (AI-system: Se/Sp 0.92/0.84; rater: 0.97/0.77; AUC 0.91 vs 0.87). In D3 decision strategies, rater-only prioritised specificity (Se/Sp 0.60/0.85), AI-system-only prioritised sensitivity (0.75/0.77), OR reduced false negatives (0.84/0.73), and AND reduced false positives (0.51/0.89), with AND yielding the lowest clinical loss at higher false-positive penalties. Conclusions The AI-system provides complementary information that becomes clinically relevant when integrated into structured human-AI-system decision rules. Context-dependent use may support minimally invasive caries management. Clinical Significance Combining human assessment with AI-system may improve preventive and operative decision-making by balancing false negatives and false positives.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ganß et al. (2026) studied this question.

synapsesocial.com/papers/69aa70e7531e4c4a9ff5b158https://doi.org/10.48620/95922
Ask AI
Helpful
Bookmark
Share
View Full Paper