PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 18, 202247 citations

MLTR: Multi-Label Classification with Transformer

View Full Paper
CXCheng XingHLHezheng LinXWXiangyu Wu

Key Points

Key points are not available for this paper at this time.

Abstract

The task of multi-label image classification is to recognize all the object labels presented in an image. Though advancing for years, small objects, and objects with high conditional probability are still the main bottlenecks of previous convolutional neural network (CNN) based models, limited by convolutional kernels' representational capacity. Recent vision transformer networks utilize the self-attention mechanism to extract the feature of pixel granularity. It expresses richer local semantic information, while insufficient for mining global spatial dependence. In this paper, we point out the three crucial problems that CNN-based methods encounter and explore the possibility of conducting specific transformer modules to settle them. We put forward a Multi-label Transformer architecture (MlTr) constructed with windows partitioning, in-window pixel attention, cross-window attention, particularly improving the performance of multi-label image classification tasks. The proposed MlTr shows state-of-the-art results on various prevalent multi-label datasets such as MS-COCO, Pascal-VOC, NUS-WIDE with 88.8%, 95.8%, 65.5% respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xing et al. (2022) studied this question.

synapsesocial.com/papers/6a15f49832de3075b8524b02https://doi.org/10.1109/icme52920.2022.9860016
Ask AI
Helpful
Bookmark
Share
View Full Paper