Unverified paper record
The effect of train-test data splitting ratio on the accuracy and reliability of yield prediction from multispectral UAV imagery: Effective reduction of training data with algorithm selection for single- and multi-date models
27 Feb 2026 · 10.21203/rs.3.rs-8762919/v1
Abstract
Abstract Context The amount and ratio of data required for training spectral grain yield (GY) prediction models remain critical for sampling reference data, but have been rarely addressed so far. Given the high cost for collecting ground truth data, common data splitting ratios which use the majority of the data for training and the rest for testing the models, should be reduced, which in turn may affect prediction accuracy. Methods Therefore, this study evaluated 11 different train-test data splitting ratios (TSR), ranging from using 5% to 95% of the data for training, and the remaining portion as test set. Models with six different machine learning algorithms were compared in winter wheat breeding trials conducted with each several thousand plots in two locations in Germany over a period of four years. The input data consisted of the UAV-based NDRE index data from individual measurements dates as well as incremental date combinations. Results The results indicate that GY prediction was relatively stable when decreasing TSR to about 0.30. Conversely, multi-date models tended to profit more from higher TSR than single-date models. Support vector machine and random forest algorithms showed relative advantage for higher TSR and multi-date models, whereas partial least squares and ridge regression were the best algorithms for lowest TSR-values. Prediction results from repeated data splitting indicates minimum R²-variability at TSR-values of 0.30 but substantially increasing R²-variation for high TSR-values. Conclusions It is concluded that decreasing TSR while considering algorithm selection can reduce costs without compromising prediction accuracy, therefore making spectral phenotyping methods more accessible and ready-to-use.
Plant phenotyping relevance
UAV multispectral画像と機械学習による収量推定について、データ分割比・アルゴリズム・予測安定性を体系的に評価しており、スペクトル表現型計測ワークフローの技術的検証が中心である。
abstractTherefore, this study evaluated 11 different train-test data splitting ratios (TSR), ranging from using 5% to 95% of the data for training, and the remaining portion as test set.
abstractPrediction results from repeated data splitting indicates minimum R²-variability at TSR-values of 0.30 but substantially increasing R²-variation for high TSR-values.
abstracttherefore making spectral phenotyping methods more accessible and ready-to-use.
Code and data availability
The preprint describes UAV multispectral NDRE data, grain yield reference data, and an R/caret modeling pipeline, but contains no public data or code deposit, no availability statement, and no author-provided URL for datasets, images, scripts, or models. The supplementary file (SupplementaryMaterial.docx) is listed but
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.