← Papers

Unverified paper record

Boosting crop disease recognition via automated image description generation and multimodal fusion

Computers and Electronics in Agriculture. · 1 Dec 2025

Abstract

In phytopathological diagnostics, traditional unimodal, vision-based methodologies often encounter limitations in performance due to the ambiguous manifestation of disease symptoms and the complexity of background environments. In particular, the robustness and generalizability of such methods are further limited under field conditions characterized by variable illumination, occlusion, and background clutter. Conversely, textual modality exhibits robust semantic encoding capabilities, enabling precise characterization of disease phenotypes and enhancing the discriminative capability of the model. However, textual modality remains underexploited and insufficiently integrated within current research paradigms. In this study, we propose a multimodal diagnostic framework that leverages large multimodal models to automatically generate structured textual descriptions from crop imagery, thereby reducing reliance on manual annotations. To further enhance cross-modal integration, a Projected Visual–Textual Discriminant (PVD) module is introduced. Empirical results indicate that by incorporating generated textual data, most multimodal architectures consistently outperform their unimodal equivalents across diverse visual backbone networks. Notably, the combination of CogAgent with CLIP (ViT-L/14) and the PVD module achieves an F1 score of 70.76%, while LLaVA combined with ResNet50+LSTM achieves 66.38%. The proposed methodology offers a practical and scalable solution for in-field crop disease diagnosis by better aligning visual and textual representations, enhancing classification performance, and obviating the need for manually curated textual annotations by automatically generating descriptions using the Automated Image Description Generation module.

Plant phenotyping relevance

作物画像から病徴を推定するマルチモーダル診断フレームワークを開発し、分類性能を比較評価しており、植物病害状態の表現型取得・推定が中心である。

abstractwe propose a multimodal diagnostic framework that leverages large multimodal models to automatically generate structured textual descriptions from crop imagery
abstractEmpirical results indicate that by incorporating generated textual data, most multimodal architectures consistently outperform their unimodal equivalents across diverse visual backbone networks.

Code and data availability

公開状態または取得可能な本文経路を確認できませんでした。

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.