The soybean dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3 (accessed on 1 April 2025).
Open resource ↗Dryad · 10.5061/dryad.41ns1rnj3 · pdf-page:12 lines:1-58Unverified paper record
Cross-Modal Data Fusion via Vision-Language Model for Crop Disease Recognition.
Sensors · 30 Jun 2025 · 10.3390/s25134096
Abstract
Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate disease identification is crucial for improving crop yield and quality. While most existing deep learning-based methods focus primarily on image datasets for disease recognition, they often overlook the complementary role of textual features in enhancing visual understanding. To address this problem, we proposed a cross-modal data fusion via a vision-language model for crop disease recognition. Our approach leverages the Zhipu.ai multi-model to generate comprehensive textual descriptions of crop leaf diseases, including global description, local lesion description, and color-texture description. These descriptions are encoded into feature vectors, while an image encoder extracts image features. A cross-attention mechanism then iteratively fuses multimodal features across multiple layers, and a classification prediction module generates classification probabilities. Extensive experiments on the Soybean Disease, AI Challenge 2018, and PlantVillage datasets demonstrate that our method outperforms state-of-the-art image-only approaches with higher accuracy and fewer parameters. Specifically, with only 1.14M model parameters, our model achieves a 98.74%, 87.64% and 99.08% recognition accuracy on the three datasets, respectively. The results highlight the effectiveness of cross-modal learning in leveraging both visual and textual cues for precise and efficient disease recognition, offering a scalable solution for crop disease recognition.
Plant phenotyping relevance
作物葉の病徴を画像・テキストから認識するマルチモーダル手法の開発と評価が研究の中心であり、植物の病害状態を直接推定している。
abstractwe proposed a cross-modal data fusion via a vision-language model for crop disease recognition.
abstractExtensive experiments on the Soybean Disease, AI Challenge 2018, and PlantVillage datasets demonstrate that our method outperforms state-of-the-art image-only approaches with higher accuracy and fewer parameters.
Code and data availability
The paper's Data Availability Statement explicitly links the public image datasets used for its crop disease recognition experiments: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). These are the phenotyping image inputs directly used in this study. The AI Challenge 2018 dataset is also公开
plantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-
Open resource ↗Kaggle · pdf-page:12 lines:1-58This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.