Unverified paper record
Structured Multi-Kernel Heteroscedastic Gaussian Process for Crop Straw-to-Grain Ratio Prediction and Uncertainty Quantification
Agronomy · 9 Aug 2026 · 10.3390/agronomy16161524
Abstract
Crop straw-to-grain ratio (SGR) estimation underpins regional straw resource assessment, yet national inventories rely on fixed coefficients that ignore structured variation across variety, environment, and phenotype. We introduce a Structured Multi-Kernel Heteroscedastic Gaussian Process (GP) framework that models SGR variation through three additive kernels heuristically motivated by the genotype–environment–phenotype (G+E+P) framework—capturing variety-associated variation, spatially structured variation, and environmental and management covariates—and employs an input-dependent noise model for prediction-specific uncertainty quantification. To prevent information leakage, target encoding and feature scaling are recomputed within each cross-validation fold. Evaluated via internal leave-one-out cross-validation on 80 rice samples (42 varieties, six Chinese provinces), the model achieves R2=0.541 with a prediction interval coverage probability of 0.95. Ablation identifies variety-associated variation as the largest contributor among the modeled factors (ΔR2=−0.024) and the multi-kernel design, by incorporating variety-specific information, substantially improves upon a covariate-only RBF GP (ΔR2=0.103). On point-prediction accuracy, Gradient Boosting achieves R2=0.58, slightly ahead of the Heteroscedastic GP (R2=0.54), underscoring that the primary advantage of the GP lies in its input-dependent uncertainty quantification. However, leave-one-county-out validation yields R2≈0 (with σ escalating to 24.4), confirming that the model does not yet generalize to unsampled counties; all reported performance is therefore internal to the nine sampled counties. The framework couples an agronomically motivated additive kernel structure with input-dependent uncertainty quantification, offering a path toward uncertainty-aware prediction from small field datasets.
Plant phenotyping relevance
作物のわら・穀粒比という植物関連形質を対象に、不確実性定量化を備えた予測手法を開発・検証しており、単なるルーチン測定ではなく計算的な形質推定が中心である。
abstractWe introduce a Structured Multi-Kernel Heteroscedastic Gaussian Process (GP) framework that models SGR variation
abstractthe primary advantage of the GP lies in its input-dependent uncertainty quantification
abstractEvaluated via internal leave-one-out cross-validation on 80 rice samples
Code and data availability
The paper's 80 field-collected rice SGR samples, phenotype measurements, and GP analysis code are described but no public dataset, image, code repository, or supplement with such assets is mentioned. NASA POWER and SoilGrids are generic external environmental data sources, not paper-specific phenotyping assets.
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.