Computational evaluation demonstrates enhanced scene classification accuracy in micro-videos, highlighting the benefit of disentangling modality noise and dynamic fusion.
Key Points
To develop a robust micro-video scene classification framework that effectively filters real-world noise and adaptively balances inconsistent multimodal signals.
Constructed a Generative Feature Disentanglement module that decomposes each modality into modality-specific, modality-shared, and noise representations by minimizing generative loss.
Designed a Dynamic Multimodal Fusion mechanism that adaptively weights and integrates specific and shared components based on their information entropy contribution.
Demonstrated qualitative and quantitative improvements in micro-video scene classification accuracy across complex, real-world video environments.
Effectively isolated redundant noise from shared and modality-specific semantics, resolving fusion weight uncertainties and semantic discrepancies.