Randomized trial demonstrates improved recommendation accuracy in multimedia recommenders, suggesting enhanced user preference modeling.
Multimedia recommenders can use behavioral records together with visual and textual item information, but unreliable interactions and sparse histories still make user preference modeling difficult. Most graph-based methods propagate messages over observed user–item edges as if all interactions were equally informative, so incidental or semantically inconsistent behaviors may distort the learned representations. The standard recommendation loss also provides limited context for modeling dependencies within a user’s historical sequence. We propose MGDSL, a MGDSL applies a multimodal-aware topology denoising module to calculate edge reliability weights for historical interactions from collaborative, textual, and visual evidence, and uses these weights for reliability-aware historical aggregation. In parallel, a masked self-supervised auxiliary task reconstructs masked items from sequence context, adding supervision for latent preference learning. Experiments on three benchmark datasets show that MGDSL consistently improves recommendation accuracy over competitive baselines, with particularly clear gains on the sparsest dataset.
No takes yet. Share an insight, caveat, or question.
Xu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: