← Papers

Unverified paper record

Machine Learning-Based Prediction of Soybean Plant Height from Agronomic Traits Across Sequential Harvests

AgriEngineering · 2 Dec 2025 · 10.3390/agriengineering7120408

Abstract

The accurate prediction of plant height is crucial for optimizing soybean cultivar selection and improving yield estimations. In this study, we investigate the potential of machine learning (ML) algorithms to predict soybean plant height (PH) based on a diverse set of agronomic parameters analyzed from forty soybean cultivars evaluated across sequential harvests. Using a comprehensive dataset, the models Elastic Net (EN), Extra Trees (ET), Gaussian Process Regressor (GPR), K-Nearest Neighbors, and XGBoost (XGB) were compared in terms of predictive accuracy, uncertainty, and robustness. Our results demonstrate that ET outperformed other models with an average correlation coefficient of 0.674, R2 of 0.426 and the lowest RMSE of 6.859 cm and MAE of 5.361 cm, while also showing the lowest uncertainty (5.07%). The proposed ML framework includes an extensive model evaluation pipeline that incorporates the Performance Index (PI), ANOVA, and feature importance analysis, providing a multidimensional perspective on model behavior. The most influential features for PH prediction were the number of stems (NS) and insertion of the first pod (IFP). This research highlights the viability of integrating explainable ML techniques into agricultural decision support systems, enabling data-driven strategies for cultivar evaluation and phenotypic trait forecasting.

Plant phenotyping relevance

大豆の草丈という植物形質を予測する機械学習フレームワークを提案し、複数モデルの精度・不確実性・頑健性を比較評価しており、形質推定手法が研究の中心である。

abstractThe proposed ML framework includes an extensive model evaluation pipeline that incorporates the Performance Index (PI), ANOVA, and feature importance analysis, providing a multidimensional perspective on model behavior.
abstractThis research highlights the viability of integrating explainable ML techniques into agricultural decision support systems, enabling data-driven strategies for cultivar evaluation and phenotypic trait forecasting.

Code and data availability

The paper's soybean agronomic phenotype dataset (320 samples, 40 cultivars, traits including PH, IFP, NS, GY) and the authors' ML analysis source code are paper-specific assets, but the Data Availability Statement restricts access to contacting the authors; no public deposit or URL is provided.

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.