The limits of applicability of vision-and-language models are defined by the of their training data. Tasks like vision question answering (VQA) require commonsense and factual information beyond what can be learned task-specific datasets. This paper investigates the injection of knowledge general-purpose knowledge bases (KBs) into vision-and-language. We use an auxiliary training objective that encourages the representations to align with graph embeddings of matching entities in KB. We empirically study the relevance of various KBs to multiple tasks and. The technique brings clear benefits to knowledge-demanding question tasks (OK-VQA, FVQA) by capturing semantic and relational knowledge from existing models. More surprisingly, the technique also benefits reasoning tasks (NLVR2, SNLI-VE). We perform probing experiments and that the injection of additional knowledge regularizes the space of, which improves the representation of lexical and semantic. The technique is model-agnostic and can expand the applicability any vision-and-language transformer with minimal computational overhead.
No takes yet. Share an insight, caveat, or question.
Shevchenko et al. (2021) studied this question.