ABSTRACT A wireless capsule endoscope (WCE) is currently the first‐line examination to investigate digestive tract diseases. Although the procedure is reliable, a large number of images are produced by the built‐in camera, usually resulting in an unmanageable amount of data that requires a significant amount of time for manual examination by clinical experts. Hence, in this paper, two new models, Mobile3ViTL and Mobile3ViTS, for high‐ and low‐resource use scenarios are proposed for potential application in gastrointestinal image recognition. The proposed models draw lessons from the transformer and include ViTBlocks, which combine the advantages of convolutional neural networks (CNNs) and the visual transformer model, and make full use of the local and global information of images. The proposed models were exhaustively compared with recent models available in the literature. Using the same training environment, the proposed models achieved the best results in terms of accuracy (98. 58% Mobile3ViTL and 98. 57% Mobile3ViTS), reduced training time (average time of 6. 5 h Mobile3ViTL and 4. 17 h Mobile3ViTS), and reduced number of parameters (approximately 46% Mobile3ViTL and 69% Mobile3ViTS compared with MobileNetV3S). In addition, in‐depth simulations were performed to verify the feasibility of the proposed models for gastrointestinal image recognition. The results achieved were found to be promising in terms of recognition accuracy, with a significant reduction in the overall training time and the number of parameters.
Ye et al. (2026) studied this question.