PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 25, 20245 citationsOpen Access

Visual CoT: Unleashing Chain-of-Thought Reasoning in Multi-Modal Language Models

View Full Paper
HSHao ShaoSQShengju QianHXHan Xiao

Key Points

  • Visual chain-of-thought reasoning improves multi-modal large language models by dynamically focusing on visual inputs and generating interpretable thoughts.
  • A new dataset comprising 373k question-answer pairs provides intermediate bounding boxes across complex visual inputs to benchmark local region identification.
  • Benchmark evaluations demonstrate that multi-turn processing with visual chain-of-thought reasoning enhances inference strategies for multi-modal language models.

Abstract

This paper presents Visual CoT, a novel pipeline that leverages the reasoning capabilities of multi-modal large language models (MLLMs) by incorporating visual Chain-of-Thought (CoT) reasoning. While MLLMs have shown promise in various visual tasks, they often lack interpretability and struggle with complex visual inputs. To address these challenges, we propose a multi-turn processing pipeline that dynamically focuses on visual inputs and provides interpretable thoughts. We collect and introduce the Visual CoT dataset comprising 373k question-answer pairs, annotated with intermediate bounding boxes highlighting key regions essential for answering the questions. Importantly, the introduced benchmark is capable of evaluating MLLMs in scenarios requiring specific local region identification. Extensive experiments demonstrate the effectiveness of our framework and shed light on better inference strategies. The Visual CoT dataset, benchmark, and pre-trained models are available to foster further research in this direction.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shao et al. (2024) studied this question.

synapsesocial.com/papers/68e72652b6db6435876a0772https://doi.org/10.48550/arxiv.2403.16999
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought2024
  2. 2Imagine while Reasoning in Space: Multimodal Visualization-of-Thought2025
  3. 3Progressive Visual Rationale for Multimodal Chain-of-Thought in Large Vision-Language Models2026
  4. 4VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models2024
  5. 5Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models2024 · 9 citations