PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 18, 20242 citationsOpen Access

Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

View Full Paper
ZXZhiyang XuNanjing Medical UniversityCFChao FengQilu University of TechnologyRSRulin ShaoUniversity of Washington

Key Points

Key points are not available for this paper at this time.

Abstract

Despite vision-language models' (VLMs) remarkable capabilities as versatile visual assistants, two substantial challenges persist within the existing VLM frameworks: (1) lacking task diversity in pretraining and visual instruction tuning, and (2) annotation error and bias in GPT-4 synthesized instruction tuning data. Both challenges lead to issues such as poor generalizability, hallucination, and catastrophic forgetting. To address these challenges, we construct Vision-Flan, the most diverse publicly available visual instruction tuning dataset to date, comprising 187 diverse tasks and 1,664,261 instances sourced from academic datasets, and each task is accompanied by an expert-written instruction. In addition, we propose a two-stage instruction tuning framework, in which VLMs are firstly finetuned on Vision-Flan and further tuned on GPT-4 synthesized data. We find this two-stage tuning framework significantly outperforms the traditional single-stage visual instruction tuning framework and achieves the state-of-the-art performance across a wide range of multi-modal evaluation benchmarks. Finally, we conduct in-depth analyses to understand visual instruction tuning and our findings reveal that: (1) GPT-4 synthesized data does not substantially enhance VLMs' capabilities but rather modulates the model's responses to human-preferred formats; (2) A minimal quantity (e.g., 1,000) of GPT-4 synthesized data can effectively align VLM responses with human-preference; (3) Visual instruction tuning mainly helps large-language models (LLMs) to understand visual features.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2024) studied this question.

synapsesocial.com/papers/68e78b93b6db6435876fd8dahttps://doi.org/10.48550/arxiv.2402.11690
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Rethinking Overlooked Aspects in Vision-Language Models2024
  2. 2Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection2024
  3. 3Generative Visual Instruction Tuning2024
  4. 4VIGC: Visual Instruction Generation and Correction2024 · 26 citations
  5. 5Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models2024 · 2 citations