PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 25, 2026Journal of KIISE0 citations

Robust Visual Question Answering Model with Data Augmentation using Siamese Network

View Full Paper
IOInJae OhYLYuEun LeeJKJung-Uk Kim

Key Points

  • The aim is to develop a robust visual question answering model that maintains performance against data alterations.
  • Designed a siamese network-based teacher-student model for visual question answering.
  • Applied knowledge distillation techniques to improve learning from augmented images.
  • Evaluated performance against Gaussian blur-altered images compared to existing models.
  • Demonstrated superior performance on Gaussian blur-altered images compared to traditional VQA models.
  • Maintained high performance by adjusting KL divergence loss, retaining knowledge from original images.

Abstract

본 연구는 Gaussian blur와 같은 데이터 변형에도 강건한 성능을 유지할 수 있는 시각적 질의 응답 (Visual Question Answering, VQA) 모델을 제안한다. 기존 VQA 모델은 VQAv2 등의 시각적인 정보에 변형이 없는 데이터셋을 중심으로 학습되어, 변형된 이미지 입력에 대해 성능이 저하되는 한계가 있었다. 이를 극복하기 위해, 본 연구에서는 Siamese 구조를 기반으로 한 teacher-student 모델을 설계하여 Knowledge Distillation (KD) 기법을 적용하였다. 제안된 방법들은 Gaussian blur가 적용된 이미지에 대해 기존 모델 대비 우수한 성능을 보였으며, 특히 KL divergence loss 비중을 선형적으로 조정하여 변형 이미지에 대한 지식을 습득하는 과정에서 원본 이미지에 대한 지식을 망각하지 않고 높은 성능을 유지할 수 있었다.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Oh et al. (2026) studied this question.

synapsesocial.com/papers/69ec598788ba6daa22dab5b8https://doi.org/10.5626/jok.2026.53.4.288
Ask AI
Helpful
Bookmark
Share
View Full Paper