Key points are not available for this paper at this time.
This systematic literature review, guided by Kitchenham and Charters and following PRISMA 2020, analyzes explainable artificial intelligence (XAI) approaches for multiclass classification models, with an emphasis on explaining class differentiation and the relationship between feature contributions and changes in prediction probabilities. The protocol was defined in advance, but it was not preregistered. Searches were conducted in Scopus, Web of Science, SpringerLink, and ScienceDirect (2020–2025) using PICOC-based strings and explicit eligibility criteria. Following the PRISMA flow, 108 studies were included out of 8697 identified records. The most frequently reported approaches are based on feature contribution/attribution (e.g., SHAP, LIME, CAM, and Grad-CAM) and counterfactual explanations, with prominent applications in medicine, finance, and cybersecurity. Although several works analyze local contributions and, separately, probability variations, the synthesis reveals a methodological gap: there is a lack of a formal and explicit instance-level framework that quantitatively connects the differential contribution of a feature (e.g., SHAP values) with the probability variation between classes to explain class differentiation. In practical terms, such a linkage enables instance-level justification of why a model favors class A over a competing class B, improving traceability and decision support in high-stakes settings (e.g., differential diagnosis and risk assessment). These findings point to future directions toward more rigorous comparative local explanations in multiclass settings.
Romero et al. (Thu,) studied this question.