PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 1, 2026ENLIGHTEN (Jurnal Bimbingan dan Konseling Islam)0 citationsOpen Access

Attention-Enhanced Cross-Modality Alignment for Adapting Vision-Language Models

SXShaoying XueXYXiaochen YangXLXiaoxu Li

Key Points

Key points are not available for this paper at this time.

Abstract

Prompt learning is an effective approach for adapting pre-trained vision-language models (VLMs) to a variety of downstream tasks. However, prompts designed manually or generated by large language models may not effectively capture key discriminative visual features. In addition, pre-trained VLMs may not align images and text well at a fine-grained level. To address these two issues, we propose an attention-enhanced cross-modality alignment network, which includes an adaptive channel attention (ACA) module and a cross-modal measurement (CMM) module. The ACA module adapts the existing efficient channel attention to highlight discriminative visual and textual features. The CMM module leverages four pairs of image-text similarities across both frozen and learnable branches, improving the alignment of fine-grained discriminative visual and textual features. Experiments show that the proposed method outperforms state-of-the-art methods on two representative tasks: base-to-novel generalization and cross-dataset evaluation. Our code is available at https: //github. com/xueshaoying/XSYAECA. git.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xue et al. (2026) studied this question.

synapsesocial.com/papers/6a08f5e35405cc787b9d06e9https://doi.org/10.1007/s10994-025-06952-5
Ask AI
Helpful
Bookmark
Share
View Full Paper