Key points are not available for this paper at this time.
Visual Question Answering (VQA) is one of the attractive topics in the field of multimedia, affective, and empathic computing to garner user interest. Unlike existing models which aim at addressing chal- lenges of VQA for the scene images, this work aims at developing a new model for Personality Traits Question Answering (PQA). It uses Twitter account information, which includes shared images, pro- file pictures, banners, text in the images, and descriptions of the images. Motivated by the accomplish- ments of the transformer, for encoding visual features of the images, a new InfoGain Multi-Axial Wavelet Vision Transformer (IgMaWaViT) is explored here. For encoding textual features in the im- ages and descriptions, a new Information Gain BERT (InfoBert) method is introduced, which can handle the variable length encoding of text by choosing the optimal discriminator. Furthermore, the model fuses encodings of images and text according to the questions on different personality traits for question answering. The model is called InfoGain Multi-Axial Wavelet Vision Transformer for Per- sonality Traits Question Answering (IgMaWaViT-PQA). To validate the efficacy of the proposed model, a dataset has been constructed, and it is used along with standard datasets for experimentation.
Wenbin Wang (Tue,) studied this question.