PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 20244 citations

Visual Question Answering

View Full Paper
ANAhmed NadaMCMin Chen

Key Points

Key points are not available for this paper at this time.

Abstract

Visual question answering (VQA) is an artificial intelligence (AI) and computer vision (CV) comprehensive task to answer questions about the visual content of an image, such as "what color is the bus?" or "how many people are in the photo?" VQA has shown great potential and importance in various domains, ranging from medical imaging applications, autonomous driving, to virtual assistants and search engines. This study develops a framework to tackle VQA research challenges by adopting and extending recent breakthroughs in attention techniques, natural language processing, and image classification models. In addition, different from other previous work that uses static question embedding, we investigate how alternative dynamic embedding models enhance the effectiveness of VQA task. The work is evaluated using the latest developed VQA v2 dataset with a 9% improvement over the results obtained with static word embedding. We also deployed the model as a cloud based VQA system to facilitate VQA tasks in real-life applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nada et al. (2024) studied this question.

synapsesocial.com/papers/68e78a66b6db6435876fd210https://doi.org/10.1109/icnc59896.2024.10556344
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Going deeper with convolutions2015 · 47,238 citations
  2. 2Proceedings of the 26th International Conference on Neural Information Processing Systems2013 · 1,328 citations
  3. 3Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019 · 8,346 citations