Heterogeneous multi-modal data that mixes categorical demographics, continuous biomarkers and high-dimensional binary indicators remains a long-standing challenge for deep learning, particularly when fusion across modalities is delegated to undifferentiated attention blocks and when the resulting predictor must be deployed under strict interpretability constraints. We propose a multi-modal mixture-of-experts deep network that injects modality-specific inductive bias into the fusion pipeline rather than treating heterogeneous streams as exchangeable token sequences. The model couples three dedicated encoders—a clinical encoder equipped with a learnable periodic embedding for the continuous variables, a dynamics encoder that re-orders binary indicators along a pseudo-temporal axis and processes them with stacked dilated temporal convolutions modulated by Feature-wise Linear Modulation, and a structural encoder that combines three Graph Attention layers with a Transformer block over a soft fully-connected graph—and integrates them through pairwise cross-pathway attention, gated fusion and a Top-2-of-6 Mixture-of-Experts router. The whole pipeline is trained end-to-end under a multi-task objective that jointly couples classification, evidential uncertainty estimation, auxiliary multi-label prediction and risk regression. To make the resulting predictor inspectable, we further design a four-layer Tree-SHAP attribution framework that operates on a distilled tree-based surrogate—trained to reproduce the deep model’s outputs rather than the raw label—and exposes the model at the global, sample, interaction and subgroup levels within a single coherent pipeline. On a programmatically simulated heterogeneous cohort of 420 cases spanning 21 multi-modal variables, the end-to-end network attains an AUC of 0.994± 0.014 0.994 ± 0.014 and an accuracy of 98.8± 1.2\,% 98.8 ± 1.2 % under 5× 5 5 × 5 repeated stratified cross-validation; an extensive analysis shows that this cohort is near-linearly separable (so conventional baselines saturate as well), that performance degrades gracefully under simulated acquisition shift and reduced class separation, that a label-permutation control collapses to chance, and that the distilled surrogate reproduces the deep predictor (probability R²=0.77 R 2 = 0.77 , class agreement $$94\,%$$ 94 % ), which makes the subsequent attribution layer faithful by construction. The hierarchical attribution analysis further shows that eight features account for over $$80\,%$$ 80 % of the predictive signal, cross-modality interactions contribute $$55.6\,%$$ 55.6 % of the top-50 interaction strength, and the per-subgroup evidence chains remain strictly symmetric around the population expectation. Together, these results—reported with their limitations made explicit—establish PD-GliomaNet and its hierarchical Tree-SHAP companion as a transferable blueprint for accurate, uncertainty-aware and inspectable multi-modal deep learning on heterogeneous data.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: