PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2024438 citationsOpen Access

Hierarchical Question-Image Co-Attention for Visual Question Answering

View Full Paper
JLJiasen LuJYJianwei YangDBDhruv Batra

Key Points

Key points are not available for this paper at this time.

Abstract

A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant answering the question. In this paper, we argue that in addition modeling where look or visual attention, it is equally important model what words listen to or question attention. We present a novel co-attention model for VQA that jointly reasons about image and question attention. In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN). Our model improves the state-of-the-art on the VQA dataset from 60.3% 60.5%, and from 61.6% 63.3% on the COCO-QA dataset. By using ResNet, the performance is further improved 62.1% for VQA and 65.4% for COCO-QA.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lu et al. (2024) studied this question.

synapsesocial.com/papers/694c536985397a08abacbd9fhttps://doi.org/10.57702/uhdmfu1r
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Reasoning about Entailment with Neural Attention2015 · 407 citations
  2. 2Going deeper with convolutions2015 · 47,236 citations
  3. 3Convolutional Neural Network Architectures for Matching Natural Language Sentences2015 · 972 citations
  4. 4Torch7: A Matlab-like Environment for Machine Learning2011 · 1,258 citations
  5. 5Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering2016 · 763 citations