PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 20162,099 citations

Stacked Attention Networks for Image Question Answering

View Full Paper
ZYZichao YangXHXiaodong HeJGJianfeng Gao

Key Points

  • The aim is to enhance the ability to answer natural language questions based on image content using SANs.
  • Developed a multiple-layer stacked attention network for progressive reasoning.
  • Conducted experiments on four distinct image question answering datasets.
  • Utilized semantic representation of questions to locate relevant image regions.
  • SANs significantly outperformed previous state-of-the-art models in answering accuracy.
  • Visualizations showed progressive identification of relevant visual clues layer-by-layer.

Abstract

This paper presents stacked attention networks (SANs) that learn to answer natural language questions from images. SANs use semantic representation of a question as query to search for the regions in an image that are related to the answer. We argue that image question answering (QA) often requires multiple steps of reasoning. Thus, we develop a multiple-layer SAN in which we query an image multiple times to infer the answer progressively. Experiments conducted on four image QA data sets demonstrate that the proposed SANs significantly outperform previous state-of-the-art approaches. The visualization of the attention layers illustrates the progress that the SAN locates the relevant visual clues that lead to the answer of the question layer-by-layer.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2016) studied this question.

synapsesocial.com/papers/69c5e7d20db0f1f44715019chttps://doi.org/10.1109/cvpr.2016.10
Ask AI
Helpful
Bookmark
Share
View Full Paper