← Papers

Unverified paper record

mIT-CMCA: a cross-modal category alignment framework for robust maize disease identification.

Scientific reports · 21 Apr 2026 · 10.1038/s41598-026-48962-w

Abstract

Accurate identification of maize diseases is crucial for safeguarding global food security. Traditional image-based methods often struggle with lighting variations, occlusions, and noise, limiting their robustness and generalisation. Multimodal approaches that integrate visual and textual information have shown promise. However, these methods frequently require manually curated textual descriptions for each image, increasing data collection costs and limiting scalability and practical implementation. To address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA). This approach enforces category-level alignment between image and text modalities, enabling more accurate and interpretable cross-modal mapping. First, we construct cross-modal representations by aligning image and text modalities at the category level within a shared embedding space. Second, inspired by contrastive learning, we introduce a Cross-Modal Category Alignment (CMCA) loss based on category-level textual descriptions, reducing annotation complexity. Finally, we present an Efficient Channel-Spatial Hybrid Attention (CSHA) module that preserves inter-class boundaries while incurring minimal computational overhead, thereby enhancing feature discriminability under complex conditions. Experimental results on the maize subset of the PlantVillage dataset (MPVD) show that mIT-CMCA achieves 99.48% accuracy, 99.28% precision, 99.54% recall, and 99.41% F1-score. These results represent improvements of 0.24%, 0.13%, 0.17%, and 0.15% over the strongest vision-only baseline, MaxViT_tiny. On the self-built Maize Leaf-Field dataset (MLFD), the model achieves 93.67% accuracy, 93.76% precision, 93.67% recall, and 93.71% F1-score. It uses only 8.27 million parameters, which is 71.6% fewer than MaxViT_tiny. Its model size is 32.13 MB, which is 72.3% smaller. The proposed method also outperforms comparative models in robustness experiments under artificially added perturbations. These results demonstrate that mIT-CMCA achieves a favorable balance between accuracy and efficiency, making it suitable for practical agricultural deployment.

Plant phenotyping relevance

トウモロコシ葉画像から病害状態を推定する画像・マルチモーダル手法の開発と性能評価が研究の中心であり、植物病害フェノタイピングに該当する。

abstractwe proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA).
abstractExperimental results on the maize subset of the PlantVillage dataset (MPVD) show that mIT-CMCA achieves 99.48% accuracy
abstractThe proposed method also outperforms comparative models in robustness experiments under artificially added perturbations.

Code and data availability

The paper's maize disease identification analysis code and trained models are explicitly stated as publicly available in a GitHub repository. The phenotype image datasets (MPVD subset and self-built MLFD) are not publicly available and require contacting the corresponding author.

Codepublic

Code availability The code and models are available in the GitHub repository at https://github.com/TANGFEILONG626/mIT-CMCA..

Open resource ↗TANGFEILONG626/mIT-CMCA · html-lines:673-695

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.