Model evaluation demonstrates superior accuracy across six video-language benchmarks, highlighting the efficacy of fine-grained contrastive alignment for temporal localization.
Key Points
Investigate the limitations of standard video-text pre-training architectures on localization-oriented tasks and design a framework optimized for temporal grounding.
Introduced LocVTP, integrating fine-grained clip-word contrastive alignment to complement standard coarse-grained video-text representations.
Implemented a context projection head paired with a temporal-aware contrastive loss to capture contextual and sequential relationships across frames.
Evaluated representations across four downstream tasks spanning six benchmark video datasets for both retrieval and localization performance.
Achieved state-of-the-art performance across all six evaluated datasets on both retrieval-based and localization-based tasks.
Ablation studies confirmed that the clip-word correspondence scheme and temporal-aware loss substantially improve temporal reasoning capabilities over coarse alignment alone.