← Papers

Unverified paper record

Generalizability of machine learning models for plant traits using hyperspectral reflectance data: The case of maize

bioRxiv · 6 Jul 2025 · 10.1101/2025.07.03.661070

Abstract

Hyperspectral reflectance provides rapid and precise phenotyping of plants in a non-destructive manner both in field and well-controlled settings. The resulting data have been used to devise machine learning (ML) models for paired measurements of different traits in diverse plants and crops. Yet, despite advances in using of hyperspectral data to reliably predict crop traits of interest, there are pressing issues concerning the training of ML models, the aggregation of data from crop field trials, and the generalizability of the models in different prediction settings. We collected hyperspectral reflectance data along with 25 anatomical, gas exchange, and chlorophyll fluorescence traits from 320 recombinant inbred lines of a maize Multi-Parent Advanced Generation Inter-Cross population grown across three consecutive seasons. We use these data to systematically: (1) compare the performance of representative ML models for different traits, including slow fluorescence kinetics whose predictability by hyperspectral data has not yet been investigated, (2) evaluate the ML model performance in prediction scenarios concerning unseen genotypes, unseen seasons, and the combination thereof, (3) investigate the effects of data aggregation of ML model performance. These problems are addressed in a rigorous nested cross-validation setting that provides a template for adequate assessment of performance of ML models for diverse crop traits considering the particularities of the experimental design.

Plant phenotyping relevance

トウモロコシのハイパースペクトル反射データによる形質推定について、複数の機械学習モデル、未知遺伝子型・季節への汎化性能、データ統合の影響を系統的かつネスト化交差検証で評価しており、フェノタイピング手法の検証が中心である。

abstractWe use these data to systematically: (1) compare the performance of representative ML models for different traits, including slow fluorescence kinetics whose predictability by hyperspectral data has not yet been investigated, (2) evaluate the ML model performance in prediction scenarios concerning unseen genotypes, unseen seasons, and the combination thereof, (3) investigate the effects of data aggregation of ML model performance.
abstractThese problems are addressed in a rigorous nested cross-validation setting that provides a template for adequate assessment of performance of ML models for diverse crop traits considering the particularities of the experimental design.

Code and data availability

The paper's data availability statement explicitly provides all code and raw data (hyperspectral reflectance and trait measurements) for reproducibility via the authors' public GitHub repository.

Codepublic

idge, Cambridge, UK 9 † These authors contributed equally. 10 * Corresponding authors. 11 12 Email address: 13 rudan.xu@uni-potsdam.de 14 jfergu@essex.ac.uk 15 jk417@cam.ac.uk 16 nikoloski@mpimp-golm.mpg.de 17 18 Data availability statement 19 All code and raw data to ensure reproducibility of the results can be accessed at: 20 https://github.com/Rudan-X/HyperspectralML 21 22 Funding statement: 23 J.F. was supported by the European Union’s Horizon 2020 research and innovation program 24 grant 862201 (to J.K. and Z.N.). R.X. was supported by the International Max Planck Research 25 School "Molecular Plant Science" between the Max Planck Institute of Molecular Plant 26 Physiology and the Unive

Open resource ↗Rudan-X/HyperspectralML · pdf-raw-page:1 lines:1-71

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.