PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 21, 2026Frontiers in Plant Science0 citationsOpen Access

Distilled vision transformers with CNN fusion for robust cashew apple maturity prediction

LSL. SumalathaJMJeevaratnam MudidanaPVP. M. Durai Raj Vincent

Key Points

Key points are not available for this paper at this time.

Abstract

Introduction: Cashew apple is a nutrient-rich fruit containing abundant minerals, vitamins, and energy. However, its fleshy texture and delicate skin significantly limit its storage life and market value. Accurate maturity grading is therefore essential for improving post-harvest management and transportation efficiency. Methods: This study proposes a lightweight vision transformer (ViT) student model trained using multi-granular knowledge distillation (KD) from a stronger data-efficient image transformer (DeiT)-Base teacher. The distillation framework integrates response-based soft-label supervision, attention transfer, and token-level feature regression to enhance representation learning under limited data conditions. Auxiliary lightweight architectures, including MobileNet, ConvNeXt, and EdgeNeXt, were trained independently to provide complementary predictions, and a weighted fusion strategy was employed for ensemble evaluation. Results: The proposed ensemble ViT-KD with EdgeNeXt achieved 90% accuracy under the evaluated test split. To ensure statistical reliability and address potential partition bias, a stratified fivefold cross-validation was conducted on the dataset, yielding a mean accuracy of 86.89% ± 2.89% with consistent F1 scores and recall. The relatively low variance across the folds indicates stable internal generalization. Comparative experiments with conventional convolutional neural network (CNN) baselines and lightweight CNN baselines such as MobileViT-S and ShuffleNetV2 were performed, with the proposed ensemble framework achieving improved accuracy while maintaining computational efficiency. Computational analysis indicates that the stand-alone distilled ViT maintains a real-time inference capability of 8.79 ms per image, which supports suitability for edge-oriented agricultural applications. Discussion: These results highlight the effectiveness of knowledge-distilled lightweight transformers for data-efficient maturity grading of cashew apples.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sumalatha et al. (2026) studied this question.

synapsesocial.com/papers/6a0ef5ab1c5e2d2319fa2714https://doi.org/10.3389/fpls.2026.1787609
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020 · 22,047 citations
  2. 2Deep learning in agriculture: A survey2018 · 4,987 citations
  3. 3Tomato ripeness and shelf-life prediction system using machine learning2024 · 43 citations
  4. 4Determination of Tomato Fruit Stages Using Principal Component Analysis and Fuzzy Logic Algorithm2022 · 14 citations
  5. 5Assessing deep learning model robustness for banana ripeness classification under varying illumination conditions2025 · 5 citations