Immunohistochemistry (IHC) is a crucial step in the diagnostic and therapeutic management of colorectal carcinoma (CRC). This study aimed to explore the feasibility and diagnostic accuracy of general-purpose vision-language models (VLMs), specifically Gemini (2.5 Pro) and ChatGPT-4.0, for the automated classification of MMR status using immunohistochemical whole-slide images (WSIs). A retrospective cohort of 50 CRC cases, stratified by reference MMR phenotype, was analyzed, with models performing a dual assessment: individual IHC evaluation and an overall result (Proficient Mismatch Repair, pMMR, or Deficient Mismatch Repair, dMMR). The AI models achieved 100% specificity in correctly classifying cases with retention of all four MMR proteins as pMMR (20/20 cases). However, the overall sensitivity for detecting dMMR cases was 65% (13/20 cases), with lower concordance observed for MSH2/MSH6 loss (60%) compared to MLH1/PMS2 loss (70%). This drop in sensitivity was attributed to the models' reluctance to confirm true protein loss, often categorizing the result as "Indeterminate" or "Inadequate". Crucially, the diagnostic certainty of the AI models increased dramatically only when a clear Internal Positive Control (IPC) was visible in the image, establishing the IPC as an indispensable feature for robust dMMR identification and differentiating true biological loss from technical artifacts. • General - purpose vision–language AI models show excellent specificity (100%) for identifying pMMR colorectal cancers from IHC images. • Sensitivity for detecting dMMR remains limited (65%), with frequent indeterminate results, especially in cases lacking clear internal controls. • Detection of an internal positive control is essential for reliable AI-based MMR deficiency assessment in routine pathology.
Salzano et al. (Sun,) studied this question.