Unverified paper record
MMIU-Net: an encoder–decoder architecture based on multimodal feature fusion for wheat yield prediction under drought stress
Figshare · 28 May 2026 · 10.6084/m9.figshare.32463509.v1
Abstract
To address the poor regression performance caused by strong spatiotemporal heterogeneity, inconsistent information scales and complex feature relationships in field-based multimodal data, this study proposes a hierarchical fusion framework embedded within an encoder – decoder network. The framework integrates multi-scale, interpretable, and cross-modal representations through coordinated modules that bridge feature discrepancy understanding and information flow regulation. A multi-scale parallel pathway structure is designed to enhance joint perception of local and global information by leveraging feature mappings with different receptive fields. An interpretable feature importance allocation strategy is further introduced to improve the backbone network’s ability to provide dynamic guidance on feature contributions. This enables the model to perform adaptive weighting and feature selection during multimodal fusion. In addition, a cross-modal dense interaction and gated fusion mechanism is constructed to regulate information flow and capture fine-grained associations across modalities. The improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties. Results show that, in the first year, the highest prediction accuracy reaches an R2 of 0.8112 with an rRMSE of 16.85%. Validation using data from the same planting region in the second and third years yields a highest R2 of 0.8107 and 0.7986, with corresponding rRMSE values of 17.92% and 17.27%, respectively. Compared with other deep learning models within the same year, the proposed approach improves R2 by up to 30.52%, 33.96% and 32.79% across the three years, while reducing rRMSE by up to 41.59%, 45.01% and 46.43%. The results demonstrate that the coordinated interaction among modules establishes an integrated optimization pathway that spans from feature discrepancy understanding to information flow regulation, while maintaining interpretability in the decision process. Under complex field conditions with multiple sources of uncertainty, the proposed framework achieves stable module contributions ranging from 5% to 10% based on cross-validation and t-test analyses. This effectively alleviates the difficulty of efficient multimodal feature fusion for robust yield prediction under stress conditions. The study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings.Develops a weight-guided method to enhance interpretability of multimodal features.Designs a cross-modal dense interaction and gated fusion mechanism.Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings. Develops a weight-guided method to enhance interpretability of multimodal features. Designs a cross-modal dense interaction and gated fusion mechanism. Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features.
Plant phenotyping relevance
マルチモーダル作物センシングから小麦収量を推定する encoder–decoder 手法を開発し、複数年・品種データで検証しており、表現型推定手法が研究の中心である。
abstractthis study proposes a hierarchical fusion framework embedded within an encoder – decoder network.
abstractThe study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping.
abstractThe improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties.
Code and data availability
公開論文であることは確認できましたが、現在の公式API・許可済み取得経路では本文を自動取得できませんでした。
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.