PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2025IEEE Transactions on Pattern Analysis and Machine Intelligence21 citations

Scaling up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations

View Full Paper
YZYiyuan ZhangXDXiaohan DingXYXiangyu Yue

Key Points

  • Large-kernel ConvNets achieved an ImageNet accuracy of 88.0%, enhancing model performance and scalability.
  • The proposed UniRepLKNet architecture emphasizes large convolutional kernels for better spatial information capture.
  • Observational analysis highlights that large-kernel designs surpass smaller-kernel CNNs in terms of effective receptive fields.
  • These findings support the potential for large-kernel ConvNets to serve various modalities, including video recognition.

Abstract

This paper proposes the paradigm of large convolutional kernels in designing modern Convolutional Neural Networks (ConvNets). We establish that employing a few large kernels, instead of stacking multiple smaller ones, can be a superior design strategy. Our work introduces a set of architecture design guidelines for large-kernel ConvNets that optimize their efficiency and performance. We propose the UniRepLKNet architecture, which offers systematical architecture design principles specifically crafted for large-kernel ConvNets, emphasizing their unique ability to capture extensive spatial information without deep layer stacking. This results in a model that not only surpasses its predecessors with an ImageNet accuracy of 88.0%, an ADE20K mIoU of 55.6%, and a COCO box AP of 56.4% but also demonstrates impressive scalability and performance on various modalities such as time-series forecasting, audio, point cloud, and video recognition. These results indicate the universal modeling abilities of large-kernel ConvNets with faster inference speed compared with vision transformers. Our findings reveal that large-kernel ConvNets possess larger effective receptive fields and a higher shape bias, moving away from the texture bias typical of smaller-kernel CNNs. All codes and models are publicly available at https://github.com/AILab-CVC/UniRepLKNet, promoting further research and development in the community.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68af4cd8ad7bf08b1ead613dhttps://doi.org/10.1109/tpami.2025.3600702
Ask AI
Helpful
Bookmark
Share
View Full Paper