← Papers

Unverified paper record

MMIU-Net: an encoder–decoder architecture based on multimodal feature fusion for wheat yield prediction under drought stress

International Journal of Remote Sensing · 28 May 2026 · 10.1080/01431161.2026.2675717

Abstract

To address the poor regression performance caused by strong spatiotemporal heterogeneity, inconsistent information scales and complex feature relationships in field-based multimodal data, this study proposes a hierarchical fusion framework embedded within an encoder – decoder network. The framework integrates multi-scale, interpretable, and cross-modal representations through coordinated modules that bridge feature discrepancy understanding and information flow regulation. A multi-scale parallel pathway structure is designed to enhance joint perception of local and global information by leveraging feature mappings with different receptive fields. An interpretable feature importance allocation strategy is further introduced to improve the backbone network’s ability to provide dynamic guidance on feature contributions. This enables the model to perform adaptive weighting and feature selection during multimodal fusion. In addition, a cross-modal dense interaction and gated fusion mechanism is constructed to regulate information flow and capture fine-grained associations across modalities. The improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties. Results show that, in the first year, the highest prediction accuracy reaches an R2 of 0.8112 with an rRMSE of 16.85%. Validation using data from the same planting region in the second and third years yields a highest R2 of 0.8107 and 0.7986, with corresponding rRMSE values of 17.92% and 17.27%, respectively. Compared with other deep learning models within the same year, the proposed approach improves R2 by up to 30.52%, 33.96% and 32.79% across the three years, while reducing rRMSE by up to 41.59%, 45.01% and 46.43%. The results demonstrate that the coordinated interaction among modules establishes an integrated optimization pathway that spans from feature discrepancy understanding to information flow regulation, while maintaining interpretability in the decision process. Under complex field conditions with multiple sources of uncertainty, the proposed framework achieves stable module contributions ranging from 5% to 10% based on cross-validation and t-test analyses. This effectively alleviates the difficulty of efficient multimodal feature fusion for robust yield prediction under stress conditions. The study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping.

Plant phenotyping relevance

マルチモーダル農業センシングから小麦収量という植物形質を推定するエンコーダ・デコーダ手法を開発し、複数年・品種データで検証しているため、フェノタイピング手法が中心的である。

abstractthis study proposes a hierarchical fusion framework embedded within an encoder – decoder network.
abstractThe improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties.
abstractThe study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping.

Code and data availability

公開本文の所在を確認できませんでした。非公開または購読が必要な可能性があります。

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.