PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 24, 202424 citationsOpen Access

LAMM: Label Alignment for Multi-Modal Prompt Learning

View Full Paper
JGJingsheng GaoJRJiacheng RuanSXSuncheng Xiang

Key Points

Key points are not available for this paper at this time.

Abstract

With the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named LAMM, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2. 31 (\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https: //github. com/gaojingsheng/LAMM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gao et al. (2024) studied this question.

synapsesocial.com/papers/68e72954b6db6435876a2e35https://doi.org/10.1609/aaai.v38i3.27950
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ImageNet: A large-scale hierarchical image database2009 · 63,148 citations
  2. 2Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision2021 · 1,200 citations
  3. 3MaPLe: Multi-modal Prompt Learning2023 · 851 citations
  4. 4Language-driven Semantic Segmentation2022 · 164 citations
  5. 5A Survey of Vision-Language Pre-Trained Models2022 · 43 citations