Unverified paper record
Enhancing plant disease detection through multi-modal integration of visual and textual data.
Plant methods · 24 Mar 2026 · 10.1186/s13007-026-01521-w
Abstract
Plant diseases pose a significant threat to global agriculture, impacting crop yields and quality. Early and accurate detection is essential for effective health management but remains challenging due to visual similarity among diseases and complex field backgrounds. This study introduces AgriMM, a novel multi-modal detection framework that integrates visual images with expert-validated textual descriptions to improve diagnostic precision. The framework features three key innovations: a Hybrid Convolutional-Attention Collaborative Backbone (HCACB) to capture both fine-grained lesions and global context; a Context-enhanced Visual-Language Path Aggregation Network (CVL-PAN) for multi-scale feature fusion; and an Adaptive Region-Text Contrastive Learning (AR-TCL) module to enforce precise semantic alignment. We constructed a comprehensive dataset comprising 30,000 images and detailed symptom descriptions across five major crops (tomato, cucumber, pepper, eggplant, and squash). Experimental results demonstrate that AgriMM achieves a mean Average Precision (mAP) of 95.2%, significantly outperforming state-of-the-art unimodal baselines by 11.6%. These findings confirm that integrating linguistic semantic priors effectively resolves visual ambiguity, providing a robust tool for precision agriculture and sustainable crop protection.
Plant phenotyping relevance
植物病害の症状を画像から検出・診断するマルチモーダル手法を開発し、データセットと性能比較で検証しているため、植物フェノタイピング手法が中心である。
abstractThis study introduces AgriMM, a novel multi-modal detection framework that integrates visual images with expert-validated textual descriptions to improve diagnostic precision.
abstractExperimental results demonstrate that AgriMM achieves a mean Average Precision (mAP) of 95.2%, significantly outperforming state-of-the-art unimodal baselines by 11.6%.
Code and data availability
The supplied blocks describe a self-constructed 30,000-image multi-modal disease dataset and the AgriMM model, but contain no public deposit, availability statement, URL, or code release for the dataset, images, annotations, or trained model. The only URL present is the Creative Commons license notice, which is not a论文
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.