Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 9, 2025Open Access

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

View Full Paper
Ask AI
Bookmark
Share

Authors

KYKaining YingHDHenghui DingGJGuangquan Jie

Discussion

Loading...

Member takes

Overview

Proposed OmniAVS dataset enhances multimodal expressions and integrates reasoning in RAVS for improved performance.

Key Points

  • Extensive experiments demonstrate that OISA outperforms existing methods in multimodal reasoning and segmentation.
  • The OmniAVS dataset includes 2,104 videos and 61,095 multimodal referring expressions for diverse content understanding.
  • New innovations in OmniAVS emphasize complex reasoning and world knowledge alongside standard audio-visual cues.
  • OISA leverages MLLM to facilitate improved comprehension and reasoning-based segmentation of audiovisual content.

Cite This Study

Ying et al. (2025) studied this question.

synapsesocial.com/papers/68e7f0af2d7e30942762c81ahttps://doi.org/10.48550/arxiv.2507.22886
View Full Paper
Ask AI
Bookmark
Share