PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 20250 citationsOpen Access

Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models

View Full Paper
KSKai SunYBYushi BaiZYZhen Yang

Key Points

  • The proposed framework significantly enhances geometric understanding using hard negative contrastive learning methods.
  • Experiments demonstrate that the MMGeoLM model outperforms existing models on various geometric reasoning benchmarks.
  • The methodology involves using image-based and text-based hard negatives to optimize contrastive learning efficacy.
  • Ablation studies reveal key insights into hard negative strategies critical for geometric reasoning tasks.

Abstract

Benefiting from contrastively trained visual encoders on large-scale natural scene images, Large Multimodal Models (LMMs) have achieved remarkable performance across various visual perception tasks. However, the inherent limitations of contrastive learning upon summarized descriptions fundamentally restrict the capabilities of models in meticulous reasoning, particularly in crucial scenarios of geometric problem-solving. To enhance geometric understanding, we propose a novel hard negative contrastive learning framework for the vision encoder, which combines image-based contrastive learning using generation-based hard negatives created by perturbing diagram generation code, and text-based contrastive learning using rule-based negatives derived from modified geometric descriptions and retrieval-based negatives selected based on caption similarity. We train CLIP using our hard negative learning method, namely MMCLIP (Multimodal Math CLIP), and subsequently train an LMM for geometric problem-solving. Experiments show that our trained model, MMGeoLM, significantly outperforms other open-source models on three geometric reasoning benchmarks. Even with a size of 7B, it can rival powerful closed-source models like GPT-4o. We further conduct ablation studies to analyze three key factors: hard negative types, the efficiency of image-based negatives, and training configurations. These analyses yield important insights into optimizing hard negative strategies for geometric reasoning tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sun et al. (2025) studied this question.

synapsesocial.com/papers/68dc12c58a7d58c25ebb0abbhttps://doi.org/10.48550/arxiv.2505.20152
Ask AI
Helpful
Bookmark
Share
View Full Paper