Key points are not available for this paper at this time.
Understanding what happens in surveillance videos is critical for human–machine interaction in the Internet of Things (IoT). Among key tasks, video moment localization (VML) aims to locate the moment within untrimmed videos that semantically corresponds to a given query. Current studies predominantly emphasize visual components of videos, often neglecting the rich semantic content in the accompanying audio. Moreover, these approaches typically employ fixed fusion strategies during both training and inference phases, leading to limited flexibility and adaptability due to rigid architectures. To address these challenges, we propose a novel Progressive Dynamic Interaction Network with Audio Supplement (PDIN). Specifically, this framework employs a graph fusion approach to integrate audio and visual information, while addressing potential discordance between the two modalities. The incorporation of audio data significantly enhances the utilization of visual features, particularly in scenarios where visual information may be obscured or incomplete. Additionally, a progressive dynamic interaction strategy is also designed, where each layer builds upon the outcomes of the previous layer. This approach systematically addresses semantic differences between modalities, thereby facilitating a more integrated and effective fusion process. The underlying principle of this strategy is transferable to other multimodal tasks. Extensive experiments on three benchmark datasets (ActivityNet Captions, Charades-STA and QVHighlights) demonstrate that PDIN consistently outperforms existing methods.
Wu et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: