Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
April 17, 2026ACM Transactions on Information Systems

Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog

View Full Paper
Ask AI
Bookmark
Share

Authors

HZH ZHANGHong Kong Polytechnic UniversityMLMeng LiuYFYisen FengShenzhen Institute of Information Technology

Discussion

Loading...

Member takes

Implication

Experiments demonstrate improved video-grounded dialog responses by integrating dialog history and video content.

Key Points

  • The aim is to enhance responses in video-grounded dialog by better understanding both dialog history and video content.
  • Developed an iterative search and reasoning framework.
  • Utilized a textual encoder with path search and aggregation to identify key dialog cues.
  • Implemented a visual encoder featuring an iterative reasoning network to extract visual evidence.
  • Employed a pre-trained GPT-2 model as the answer generator.
  • Framework shows improved performance on three public datasets.
  • Successfully integrates complex dialog history and video information.
  • Generates coherent responses based on visual and textual cues.

Cite This Study

ZHANG et al. (2026) studied this question.

synapsesocial.com/papers/69e1cfe05cdc762e9d858e88https://doi.org/10.1145/3808220
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning2022 · 297 citations
  2. 2Active scene recognition with vision and language2011 · 28 citations
  3. 3Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model2023 · 20 citations
  4. 4Multimodal Dialog System: Relational Graph-based Context-aware Question Understanding2021 · 31 citations
  5. 5Recursive Visual Attention in Visual Dialog2019 · 139 citations