← Papers

Unverified paper record

Explainable Machine Learning Prediction of Soybean Lodging Grade and Key Trait Analysis Under High-Density Drip Irrigation Cultivation

Agronomy · 26 Aug 2026 · 10.3390/agronomy16171633

Abstract

To establish an accurate and interpretable prediction framework for soybean lodging grade and clarify the core regulatory traits and differentiated driving mechanisms of soybean lodging under high-density drip irrigation cultivation, 356 spring soybean germplasm accessions were used as experimental materials in this study. Morphological and mechanical traits including plant height (PH), stem pulling force (SPF), internode number (IN) and petiole length (PL) were measured over two consecutive years of field phenotyping. Two composite evaluation indices, plant height/stem pulling force ratio (PH/SPF) and plant height/internode number ratio (PH/IN), were further constructed. Four machine learning algorithms were adopted to develop multi-classification models for soybean lodging grade prediction. SHAP analysis combined with three global sensitivity approaches (perturbation analysis, Sobol’ method and Morris screening) was applied to decipher the regulatory patterns of key traits. The results showed that lodging grade significantly affected soybean grain yield and explained 25–28% of the phenotypic yield variation; yield reduction tended to plateau under severe lodging. Compared with single indicators such as SPF and PL, the two derived composite indices could stably distinguish soybean accessions with different lodging grades and exhibited stronger discriminatory power. Model comparison revealed that the XGBoost model achieved optimal prediction accuracy and generalization stability for lodging grade, with a weighted F1-score of 95.34% on the test set, significantly outperforming the conventional linear model. Interpretability analysis demonstrated that the PH/IN, PH, and PH/SPF acted as the primary positive traits promoting lodging, while SPF was the sole protective trait. Driving factors of lodging presented obvious gradient heterogeneity: mild lodging was dominated by the imbalance of plant architecture ratio, whereas severe lodging was governed by the cumulative effects of PH and IN. Strong interactions existed among all measured traits. The interpretable machine learning framework established in this study can provide theoretical support and technical references for lodging-resistant germplasm screening and targeted plant architecture regulation for densely planted soybean under drip irrigation systems.

Plant phenotyping relevance

大豆の倒伏状態を形態・力学形質から機械学習で推定し、モデル性能比較と解釈性解析を行う枠組みが研究の中心であり、単なる生物学的実験の routine 測定ではない。

abstractTo establish an accurate and interpretable prediction framework for soybean lodging grade and clarify the core regulatory traits and differentiated driving mechanisms of soybean lodging
abstractFour machine learning algorithms were adopted to develop multi-classification models for soybean lodging grade prediction.
abstractModel comparison revealed that the XGBoost model achieved optimal prediction accuracy and generalization stability for lodging grade
abstractThe interpretable machine learning framework established in this study can provide theoretical support and technical references for lodging-resistant germplasm screening

Code and data availability

The paper's phenotypic measurements (356 soybean accessions, two-year field traits) and analysis code are not publicly deposited. The Data Availability Statement only offers contact with the corresponding author, and the Hiplot web platform is cited merely as a generic visualization tool, not as a paper-specific asset.

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.