PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 2026Scientific Reports0 citationsOpen Access

Enhancing crop disease recognition framework via vision-language model with cross-attention and gated fusion

WLW Y LiuXuzhou Medical CollegeHWHan WangNantong UniversityGWGuoqing WuNantong University

Key Points

  • The research aims to improve crop disease detection by integrating visual and textual information through a novel framework.
  • Proposed a Cross-Model fusion framework utilizing a vision-language model.
  • Employed ShuffleNet-v2 for visual feature extraction and Zhipu.ai for generating textual descriptions.
  • Used cross-attention and gated fusion mechanisms for aligning and fusing visual and textual data.
  • Model achieved 99.04% accuracy on the Soybean Disease dataset and 99.12% on the PlantVillage dataset.
  • Surpassed ShuffleNet-V2 by 1.09% and 2.53% in recognition accuracy, respectively.
  • Highlights the effective integration of visual and textual cues for accurate disease recognition.

Abstract

Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate detection of such diseases is crucial for improving both crop yield and quality. While numerous deep learning approaches rely solely on image data for disease identification, they often overlook the complementary value of textual information in enhancing visual analysis. To address this limitation and effectively fuse features from different modalities, we propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition. Our approach utilizes the Zhipu.ai multi-modal model to generate comprehensive textual descriptions of diseased crop leaves, including global description, local lesion description, and color-texture description. These textual descriptions are then encoded into feature embeddings, while visual features are extracted using the ShuffleNet-v2 model as the image encoder. Subsequently, a cross-attention module aligns and fuses the two modalities, and a gated fusion module enables dynamic feature selection during the fusion process. Extensive evaluations on the Soybean Disease and PlantVillage datasets demonstrate that our method outperforms existing image-based models in terms of accuracy. Specifically, our model achieves recognition accuracies of 99.04% and 99.12% on the respective datasets, surpassing the ShuffleNet-V2 model by 1.09% and 2.53%, respectively. These results highlight the effectiveness of Cross-Model learning in integrating visual and textual cues for accurate and efficient disease recognition, offering a scalable solution for crop disease diagnosis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/6a0d5064f03e14405aa9c249https://doi.org/10.1038/s41598-026-53376-9
Ask AI
Helpful
Bookmark
Share
View Full Paper