PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 26, 20241 citationsOpen Access

C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning

View Full Paper
JMJi MaChina Agricultural UniversityWSWei SuoNorthwestern Polytechnical UniversityPWPeng WangNorthwestern Polytechnic University

Key Points

Key points are not available for this paper at this time.

Abstract

Vision-Language Instruction Tuning (VLIT) is a critical training phase for Large Vision-Language Models (LVLMs). With the improving capabilities of open-source LVLMs, researchers have increasingly turned to generate VLIT data by using open-source LVLMs and achieved significant progress. However, such data generation approaches are bottlenecked by the following challenges: 1) Since multi-modal models tend to be influenced by prior language knowledge, directly using LVLMs to generate VLIT data would inevitably lead to low content relevance between generated data and images. 2) To improve the ability of the models to generate VLIT data, previous methods have incorporated an additional training phase to boost the generative capacity. This process hurts the generalization of the models to unseen inputs (i.e., “exposure bias” problem). In this paper, we propose a new Content Correlated VLIT data generation via Contrastive Learning (C3L). Specifically, we design a new content relevance module which enhances the content relevance between VLIT data and images by computing Image Instruction Correspondence Scores S(I2C). Moreover, a contrastive learning module is introduced to further boost the VLIT data generation capability of the LVLMs. A large number of automatic measures on four benchmarks show the effectiveness of our method.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ma et al. (2024) studied this question.

synapsesocial.com/papers/68e5ee97b6db6435875837bchttps://doi.org/10.24963/ijcai.2024/128
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning2024
  2. 2VIGC: Visual Instruction Generation and Correction2024 · 27 citations
  3. 3Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion2025
  4. 4Visual In-Context Learning for Large Vision-Language Models2024 · 5 citations
  5. 5Concept-skill Transferability-based Data Selection for Large Vision-Language Models2024