PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 22, 202273 citationsOpen Access

Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

QHQibin HouCLCheng-Ze LuMCMing‐Ming Cheng

Key Points

Key points are not available for this paper at this time.

Abstract

This paper does not attempt to design a state-of-the-art method for visual recognition but investigates a more efficient way to make use of convolutions to encode spatial features. By comparing the design principles of the recent convolutional neural networks ConvNets) and Vision Transformers, we propose to simplify the self-attention by leveraging a convolutional modulation operation. We show that such a simple approach can better take advantage of the large kernels (>=7x7) nested in convolutional layers. We build a family of hierarchical ConvNets using the proposed convolutional modulation, termed Conv2Former. Our network is simple and easy to follow. Experiments show that our Conv2Former outperforms existent popular ConvNets and vision Transformers, like Swin Transformer and ConvNeXt in all ImageNet classification, COCO object detection and ADE20k semantic segmentation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hou et al. (2022) studied this question.

synapsesocial.com/papers/6a5c6f45a4cd7c485003145dhttps://doi.org/10.48550/arxiv.2211.11943
Ask AI
Helpful
Bookmark
Share
View Full Paper