PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 16, 2026ACM Transactions on Multimedia Computing Communications and Applications5 citations

Valley: Video Assistant with Large Language Model Enhanced Ability

View Full Paper
RLRuipu LuoZZZiwang ZhaoMYMin Yang

Key Points

  • The aim is to enhance video comprehension and instruction following using a multi-modal foundation model.
  • Developed two datasets: Valley-702k and Valley-instruct-73k for video-text tasks.
  • Used ViT-L/14 as the vision encoder to process visual information.
  • Explored three temporal modeling modules for comprehensive feature extraction.
  • Adopted a two-phase training approach, first for visual input understanding and then for joint training to improve instruction-following.
  • Valley demonstrates improved performance in video comprehension tasks compared to previous models.
  • Showcased enhanced instruction-following capabilities across diverse video scenarios.
  • Experimental results indicate a strong potential for applications in various video-based tasks.

Abstract

Large Language Models (LLMs), with remarkable conversational capabilities, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiveness in joint video-language understanding has not been extensively explored. In the paper, we introduce Valley , a multi-modal foundation model designed to enable enhanced video comprehension and instruction-following capabilities. To this end, we construct two datasets, namely ‘ Valley-702k ’ and ‘ Valley-instruct-73k ’, to cover a diverse range of video-text alignment and video-based instruction tasks, such as multi-shot captions, long video descriptions, action recognition, causal inference, etc. Then, we adopt ViT-L/14 as the vision encoder and explore three different temporal modeling modules to learn multifaceted features for enhanced video understanding. In addition, we implement a two-phase training approach for Valley: the first phase focuses solely on training the projection module to facilitate the LLM's capacity to understand visual input, and the second phase jointly trains the projection module and the LLM to improve their instruction following ability. Extensive experiments demonstrate that Valley has the potential to serve as an effective video assistant, simplifying complex video understanding tasks. Our code and data are publicly available at https://github.com/RupertLuo/Valley .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Luo et al. (2026) studied this question.

synapsesocial.com/papers/6992b3b19b75e639e9b08693https://doi.org/10.1145/3796716
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ImageNet: A large-scale hierarchical image database2009 · 63,148 citations
  2. 2ActivityNet: A large-scale video benchmark for human activity understanding2015 · 2,680 citations
  3. 3GRiT: A Generative Region-to-Text Transformer for Object Understanding2024 · 22 citations
  4. 4MSR-VTT: A Large Video Description Dataset for Bridging Video and Language2016 · 1,808 citations
  5. 5HViT: Hybrid vision inspired transformer for the assessment of carotid artery plaque by addressing the cross-modality domain adaptation problem in MRI2023 · 34 citations