Depression detection using multimodal artificial intelligence (AI) is rapidly evolving, transitioning from traditional machine learning techniques to deep learning architectures built on pre-trained transformer encoders. These systems integrate textual, acoustic, visual, and physiological signal streams through early, intermediate, late, and hybrid fusion strategies to improve predictive performance. However, a significant gap persists in model interpretability: post-hoc attribution frameworks (such as SHAP and LIME) dominate recent work, accounting for over 70% of studies published in 2024–2025, while clinician-centered validation is reported in only about 18% of the full reviewed corpus. Consequently, existing explanations are rarely validated for explainability quality dimensions; for example, SHAP attributions are known to degrade under inter-modality collinearity, and reported systems frequently lack alignment with established diagnostic criteria (e.g., DSM-based or PHQ-9 symptom structures). To close the gap between predictive accuracy and clinical utility, this review jointly synthesizes multimodal datasets, feature extraction strategies, machine learning and deep learning taxonomies, fusion strategies, and explainability techniques across 72 included studies. Finally, we outline critical future directions, emphasizing standardized XAI benchmarks, clinician-in-the-loop validation studies, and knowledge-driven neuro-symbolic reasoning.
No takes yet. Share an insight, caveat, or question.
Arora et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: