PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026ACM Transactions on Multimedia Computing Communications and Applications0 citations

Resilient Semantic Pseudo-Text Embedding for Zero-Shot Video Moment Retrieval

View Full Paper
DZDonglin ZhangWSWeixiang ShiXWXiao-jun Wu

Key Points

  • The aim is to improve video moment retrieval without relying on extensive human-annotated data.
  • Developed a method called Resilient Semantic Pseudo-text Modeling (RSPT).
  • Generated initial pseudo-texts by injecting random noise into visual features.
  • Learned adaptive noise weights to model correlations between pseudo-texts and visual features.
  • Introduced a quality-aware contrastive loss to ensure semantic alignment.
  • RSPT outperformed existing competitive baselines in video moment retrieval tasks.
  • Validated through extensive experiments on Charades-STA and ActivityNet-Captions.
  • Demonstrated improved alignment of semantic representations.

Abstract

With the explosive growth of video data, video moment retrieval (VMR) has attracted increasing attention due to its ability to localize semantically relevant moments in untrimmed videos. However, existing VMR approaches usually rely on annotated video-text correspondences or temporal annotations, both of which require significant human effort and are costly to scale. Even worse, the inherent subjectivity in manual labeling often introduces inconsistencies into the training data, further complicating the issue. In this paper, we investigate the problem of Zero-Shot Video Moment Retrieval (ZS-VMR) and develop a novel method, Resilient Semantic Pseudo-text Modeling (RSPT). The core of RSPT is to construct semantically rich pseudo-text embeddings through visually guided perturbations. Specifically, RSPT first generates initial pseudo-texts by injecting random noise into visual features and then learns adaptive noise weights by modeling the correlations between these pseudo-texts and visual features. This enables the generation of diverse and semantically aligned representations from multiple perspectives. To ensure alignment with visual semantics and suppress irrelevant noise, RSPT introduces a quality-aware contrastive loss that regularizes the semantic boundaries of pseudo-texts. Extensive experiments on Charades-STA and ActivityNet-Captions show that RSPT outperforms existing competitive baselines, validating its efficacy. Code is available at https://github.com/dmcsy/RSPT.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/698d6e7b5be6419ac0d544d6https://doi.org/10.1145/3796721
Ask AI
Helpful
Bookmark
Share
View Full Paper