Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 1, 2025Open Access

Describe Anything: Detailed Localized Image and Video Captioning

View Full Paper
Ask AI
Bookmark
Share

Authors

LLLong LianYDYifan DingYGYunhao Ge

Discussion

Loading...

Member takes

Overview

This model generates accurate local descriptions in images and videos, highlighting its advances in semi-supervised learning.

Key Points

  • DAM achieves state-of-the-art performance in detailed localized captioning on seven benchmarks, improving context understanding.
  • The focal prompt enhances high-resolution encoding of targeted regions, allowing for better detail preservation.
  • The proposed SSL-based Data Pipeline expands training data from existing segmentation datasets to unlabeled web images.
  • DLC-Bench provides a new metric for evaluating localized captioning independently of reference captions.

Cite This Study

Lian et al. (2025) studied this question.

synapsesocial.com/papers/68dd91cbfe798ba2fc4987a3https://doi.org/10.48550/arxiv.2504.16072
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1DescribeEarth: Detailed localized captioning for remote sensing images2026
  2. 2Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning2024 · 23 citations
  3. 3GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration2025
  4. 4Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?2024 · 1 citations
  5. 5A Unified Visual and Linguistic Semantics Method for Enhanced Image Captioning2024