Key points are not available for this paper at this time.
Abstract Most current times Visual Question Answering (VQA) models fail to interpret the multimodal knowledge of visual and text simultaneously leading to language prior problem. Language prior or language bias is a scenario where a model is biased towards the question and provides answer favouring most frequent related answer. In this scenario, model ignores the visual features. In order to reduce language prior, researchers have proposed ensemble methods (combining multiple models each serves different purpose), balanced data methods (generating additional data), modified evaluation strategy (mostly towards giving higher penalty to biased answers during training to reduce the effects) or a modified training framework (different training approaches for biased and unbiased samples). The ensemble methods suffer reduction in performance accuracy, data balance methods generate additional bias, evaluation methods does not work as intended because of random frequency distribution of related answers, modified training framework adds complexity thus it has to segregate between biased and unbiased samples. In this paper, we propose a VQA model "Ensemble of Spatial and Channel Attention Network (ESC-Net)" to overcome language bias problem by improving the visual features. In this work, we have considered a blended approach of using both regional and global image features and using an ensemble of combined channel and spatial attention mechanism to improve visual feature. The model is a more simple and effective solution than existing methods to solve language bias. Extensive experiment shows a remarkable performance improvement of 18% on VQACP v2 dataset with a comparison to current state-of-the-art (SOTA) models.
Chowdhury et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: