PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 2019351 citations

MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment

View Full Paper
DZDa ZhangXDXiyang DaiXWXin Wang

Key Points

  • To improve natural language moment retrieval in long videos, addressing semantic and structural misalignment.
  • Developed the Moment Alignment Network (MAN) for unified moment encoding and reasoning.
  • Introduced an iterative graph adjustment network to learn temporal relations end-to-end.
  • Evaluated performance on DiDeMo and Charades-STA benchmarks.
  • MAN significantly outperformed state-of-the-art methods.
  • Achieved higher accuracy in aligning candidate moments to language queries.
  • Demonstrated superior performance on complex temporal dependencies.

Abstract

This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two crucial challenges: semantic misalignment and structural misalignment. However, existing approaches treat different moments separately and do not explicitly model complex moment-wise temporal relations. In this paper, we present Moment Alignment Network (MAN), a novel framework that unifies the candidate moment encoding and temporal structural reasoning in a single-shot feed-forward network. MAN naturally assigns candidate moment representations aligned with language semantics over different temporal locations and scales. Most importantly, we propose to explicitly model moment-wise temporal relations as a structured graph and devise an iterative graph adjustment network to jointly learn the best structure in an end-to-end manner. We evaluate the proposed approach on two challenging public benchmarks DiDeMo and Charades-STA, where our MAN significantly outperforms the state-of-the-art by a large margin.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2019) studied this question.

synapsesocial.com/papers/69dc19a8ce788f95bfb64ed9https://doi.org/10.1109/cvpr.2019.00134
Ask AI
Helpful
Bookmark
Share
View Full Paper