PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 22, 2026Natural Sciences0 citationsOpen Access

Mobile3ViT: An Improved Hybrid CNN‐Visual Transformer Model for Automatic Gastrointestinal Image Recognition

View Full Paper
BYBo YeWuhan Business UniversityGZG ZhangWuhan UniversityWZWei ZhaWuhan Business University

Key Points

  • To develop and evaluate new hybrid models for automatic gastrointestinal image recognition using deep learning techniques.
  • Proposed two models: Mobile3ViT_L and Mobile3ViT_S for high and low-resource scenarios.
  • Utilized transformer blocks combined with convolutional neural networks to exploit image information.
  • Compared performance with existing models in terms of accuracy, training time, and parameters.
  • Achieved recognition accuracy of 98.58% for Mobile3ViT_L and 98.57% for Mobile3ViT_S.
  • Reduced average training time to 6.5 hours for Mobile3ViT_L and 4.17 hours for Mobile3ViT_S.
  • Decreased the number of parameters by approximately 46% (Mobile3ViT_L) and 69% (Mobile3ViT_S) compared to MobileNetV3_S.

Abstract

ABSTRACT A wireless capsule endoscope (WCE) is currently the first‐line examination to investigate digestive tract diseases. Although the procedure is reliable, a large number of images are produced by the built‐in camera, usually resulting in an unmanageable amount of data that requires a significant amount of time for manual examination by clinical experts. Hence, in this paper, two new models, Mobile3ViTL and Mobile3ViTS, for high‐ and low‐resource use scenarios are proposed for potential application in gastrointestinal image recognition. The proposed models draw lessons from the transformer and include ViTBlocks, which combine the advantages of convolutional neural networks (CNNs) and the visual transformer model, and make full use of the local and global information of images. The proposed models were exhaustively compared with recent models available in the literature. Using the same training environment, the proposed models achieved the best results in terms of accuracy (98. 58% Mobile3ViTL and 98. 57% Mobile3ViTS), reduced training time (average time of 6. 5 h Mobile3ViTL and 4. 17 h Mobile3ViTS), and reduced number of parameters (approximately 46% Mobile3ViTL and 69% Mobile3ViTS compared with MobileNetV3S). In addition, in‐depth simulations were performed to verify the feasibility of the proposed models for gastrointestinal image recognition. The results achieved were found to be promising in terms of recognition accuracy, with a significant reduction in the overall training time and the number of parameters.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ye et al. (2026) studied this question.

synapsesocial.com/papers/69e865476e0dea528dde9cf4https://doi.org/10.1002/ntls.70065
Ask AI
Helpful
Bookmark
Share
View Full Paper