PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 2026PLoS ONE3 citationsOpen Access

Hybrid lightweight vision transformers with attention mechanism for feature extraction and classification of product designs

View Full Paper
AWAbdul WahidHKHikmat Ullah KhanANAnam Naz

Key Points

  • The study aims to enhance packaging design classification using hybrid vision transformer models.
  • Developed a hybrid architecture combining convolutional neural networks and vision transformers.
  • Analyzed an image dataset of various packaging designs.
  • Compared the performance of LeViT with CNN-ResNet-50, RegNet, and ConvNeXt.
  • LeViT achieved the highest classification accuracy of 95%.
  • Outperformed traditional CNN-based models in capturing long-range relationships.
  • Demonstrated improved feature representation for packaging design analysis.

Abstract

In modern consumer markets, product packaging strongly influences customer attention and buying decisions. Attractive and informative designs help brands stand out in competitive environments. Recently, Artificial Intelligence (AI) has been widely used to support packaging evaluation, especially for design analysis, personalized user experiences, and product recommendation systems. However, traditional deep learning models, such as CNN-based ResNet-50 architectures, often fail to capture long-range relationships and global visual context. These limitations reduce their effectiveness in complex visual tasks like packaging classification. To address this issue, this study investigates the use of vision transformer-based models for packaging design analysis. We propose LeViT, an efficient hybrid architecture that combines convolutional neural networks with vision transformers. This design enables the model to learn both local visual details and global contextual features. The proposed approach improves feature representation while maintaining computational efficiency. Experiments were conducted on an image dataset of packaging designs. The performance of LeViT was compared with state-of-the-art models, including CNN-ResNet-50, RegNet, and ConvNeXt. The results show that the proposed model achieves the highest classification accuracy of 95%, outperforming all comparison methods. These findings demonstrate the effectiveness of transformer-based architectures for packaging classification. The proposed approach offers practical benefits for retail analytics, brand assessment, and marketing decision-making.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wahid et al. (2026) studied this question.

synapsesocial.com/papers/69a91d9bd6127c7a504c0813https://doi.org/10.1371/journal.pone.0343510
Ask AI
Helpful
Bookmark
Share
View Full Paper