PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026International Journal of Computer Vision2 citations

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

View Full Paper
JLJiajun LiuYWYibing WangHMHanghang Ma

Key Points

  • The aim is to develop a Video LMM, Kangaroo, to process long-context video effectively.
  • Developed a data curation system for high-quality annotations.
  • Created a large-scale dataset for vision-language pre-training.
  • Designed a curriculum training pipeline for handling long videos.
  • Kangaroo achieves state-of-the-art performance across video understanding benchmarks.
  • It excels compared to larger models and proprietary models on long video tasks.

Abstract

Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a challenging endeavor, especially for long videos. Due to insufficient access to large-scale high-quality video data and the excessive compression of visual features, current methods exhibit limitations in effectively processing long videos. In this paper, we introduce Kangaroo , a powerful Video LMM aimed at addressing these challenges. Confronted with the issue of inadequate training data, we develop a data curation system to build a large-scale dataset with high-quality annotations for vision-language pre-training and instruction tuning. In addition, we design a curriculum training pipeline with gradually increasing resolution and number of input frames to accommodate long videos. Evaluation results demonstrate that, with 8B parameters, Kangaroo achieves state-of-the-art performance across a variety of video understanding benchmarks while exhibiting competitive results on others. Particularly, on benchmarks specialized for long videos, Kangaroo excels some larger models with over 10B parameters and proprietary models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/698be001058ab1890a13ba5ehttps://doi.org/10.1007/s11263-025-02620-2
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis2025 · 103 citations
  2. 2Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models2024 · 355 citations
  3. 3TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering2017 · 474 citations
  4. 4Determining optical flow1981 · 10,062 citations
  5. 5TempCompass: Do Video LLMs Really Understand Videos?2024 · 30 citations