← Papers

Unverified paper record

Enhancing crop disease recognition framework via vision-language model with cross-attention and gated fusion.

Scientific reports · 18 May 2026 · 10.1038/s41598-026-53376-9

Abstract

Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate detection of such diseases is crucial for improving both crop yield and quality. While numerous deep learning approaches rely solely on image data for disease identification, they often overlook the complementary value of textual information in enhancing visual analysis. To address this limitation and effectively fuse features from different modalities, we propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition. Our approach utilizes the Zhipu.ai multi-modal model to generate comprehensive textual descriptions of diseased crop leaves, including global description, local lesion description, and color-texture description. These textual descriptions are then encoded into feature embeddings, while visual features are extracted using the ShuffleNet-v2 model as the image encoder. Subsequently, a cross-attention module aligns and fuses the two modalities, and a gated fusion module enables dynamic feature selection during the fusion process. Extensive evaluations on the Soybean Disease and PlantVillage datasets demonstrate that our method outperforms existing image-based models in terms of accuracy. Specifically, our model achieves recognition accuracies of 99.04% and 99.12% on the respective datasets, surpassing the ShuffleNet-V2 model by 1.09% and 2.53%, respectively. These results highlight the effectiveness of Cross-Model learning in integrating visual and textual cues for accurate and efficient disease recognition, offering a scalable solution for crop disease diagnosis.

Plant phenotyping relevance

植物葉の病徴を画像と言語情報から認識する融合フレームワークを開発し、複数データセットで既存手法と比較評価しているため、植物フェノタイピング手法が中心である。

abstractwe propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition.
abstractExtensive evaluations on the Soybean Disease and PlantVillage datasets demonstrate that our method outperforms existing image-based models in terms of accuracy.

Code and data availability

The paper's crop disease recognition experiments use two openly available image datasets, both with explicit public availability statements in the Data Availability section: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). No author analysis code, trained models, or generated text-annotait

Datasetpublic

The datasets utilized in this study are openly accessible. The soybean dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3.

Open resource ↗Dryad · 10.5061/dryad.41ns1rnj3 · html-lines:403-424
Datasetpublic

The plantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-dataset.

Open resource ↗Kaggle · plantvillage-dataset · html-lines:403-424

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.