Abstract Background In this study, an artificial intelligence (AI) model was developed that automatically assigns Mayo 0–3 scores to colonoscopy images performed at a single center. The model’s performance was compared with CRP, histological activity (Nancy score), and observer assessments. Methods In this single-center, retrospective study, 4000 colonoscopy images (approximately 1000 frames from each Mayo category) from approximately 150 ulcerative colitis patients were analyzed. An average of 30–40 frames were selected from each patient to represent different colon segments. Data were split into 80% training and 20% validation on an episode-by-episode basis. EfficientNet-B0 (ImageNet pre-trained) was used as the model. Model performance was assessed using accuracy, macro-F1, macro-AUC, confusion matrix, and class boundary separations AIₚrobₐctiveM2M3 = p¯ (Mayo2) +p¯ (Mayo3), calculated at the episode level, was compared with CRP and Nancy score using Spearman correlation. Human-AI agreement was measured by two independent endoscopists blindly evaluating a subset of samples. Results The model achieved 61. 8% accuracy on the validation set, macro-F1 = 0. 60, and macro-AUC = 0. 94. This agreement was similar to the inter-observer agreement at the Mayo 1–2 boundary. The model performed excellently in distinguishing active from inactive (AUROC (0 vs =1) = 0. 96, AUROC (=1 vs =2) = 0. 99). A significant positive correlation was observed between AI-Mayo and CRP (ρ = 0. 6, p 0. 01). The AIₚrobₐctiveM2M3 value demonstrated high discrimination power (AUROC = 0. 85) in the presence of active histology (Nancy =1). High agreement (κ = 0. 8) was found between the two observers; model predictions largely agreed with the observers’ assessments. Conclusion The AI model distinguished Mayo 0 from 3 from colonoscopy images with near-human accuracy. The model’s high agreement with biochemical and histological activity demonstrates its potential for clinical decision support in the objective and reproducible assessment of mucosal healing. Conflict of interest: Ms. Dağcı, Gizem: Large Language Model was used while creating this abstract. Yağlı, Mehmet Akif: No conflict of interest Gurbanov, Asim: No conflict of interest Agargun, Besim Fazil: No conflict of interest Telli, Pelin: No conflict of interest Mammadov, Sabuhi: No conflict of interest Genc Ulucecen, Sezen: No conflict of interest Cavus, Bilger: No conflict of interest Demir, Kadir: No conflict of interest Besisik, Fatih: No conflict of interest Kaymakoglu, Sabahattin: No conflict of interest Akyüz, Filiz: No conflict of interest
Dağcı et al. (Thu,) studied this question.