← Papers

Unverified paper record

A latent factor approach to hyperspectral time series data for multivariate genomic prediction of grain yield in wheat

arXiv (Cornell University) · 9 Jan 2026 · 10.48550/arxiv.2601.05842

Abstract

High-dimensional time series phenotypic data is becoming increasingly common within plant breeding programmes. However, analysing and integrating such data for genetic analysis and genomic prediction remains difficult. Here we show how factor analysis with Procrustes rotation on the genetic correlation matrix of hyperspectral secondary phenotype data can help in extracting relevant features for within-trial prediction. We use a subset of Centro Internacional de Mejoramiento de Maíz y Trigo (CIMMYT) elite yield wheat trial of 2014-2015, consisting of 1,033 genotypes. These were measured across three irrigation treatments at several timepoints during the season, using manned airplane flights with hyperspectral sensors capturing 62 bands in the spectrum of 385-850 nm. We perform multivariate genomic prediction using latent variables to improve within-trial genomic predictive ability (PA) of wheat grain yield within three distinct watering treatments. By integrating latent variables of the hyperspectral data in a multivariate genomic prediction model, we are able to achieve an absolute gain of .1 to .3 (on the correlation scale) in PA compared to univariate genomic prediction. Furthermore, we show which timepoints within a trial are important and how these relate to plant growth stages. This paper showcases how domain knowledge and data-driven approaches can be combined to increase PA and gain new insights from sensor data of high-throughput phenotyping platforms.

Plant phenotyping relevance

航空機搭載ハイパースペクトルセンサーによる植物表現型時系列データから潜在特徴を抽出し、収量予測に統合する解析手法が研究の中心であるため。

abstractfactor analysis with Procrustes rotation on the genetic correlation matrix of hyperspectral secondary phenotype data can help in extracting relevant features for within-trial prediction
abstractThis paper showcases how domain knowledge and data-driven approaches can be combined to increase PA and gain new insights from sensor data of high-throughput phenotyping platforms.

Code and data availability

The paper's Data and code statement provides public GitHub repositories containing the authors' analysis scripts for the hyperspectral latent-factor/Procrustes workflow and the glfBLUP R package implementing the genomic prediction methodology. The hyperspectral phenotype dataset itself is only available upon request, i

Codepublic

y of secondary trait data and successful integration in multivariate genomic prediction. As such, this method can contribute to a greater understanding of high-dimensional data in plant breeding trials. Data and code Scripts to generate the hyperspectral datasets, as well as the results presented in this paper, are available at https://github.com/KunstJF/glfBLUP-Procrustes . The glfBLUP methodology is implemented in an R-package available at https://github.com/KillianMelsen/glfBLUP . The hyperspectral dataset is available upon reasonable request from J. Crossa References Antonio et al. (2022) O. Antonio, M. López, A. Montesinos López, and J. Crossa Multivariate statistical machine learning m

Open resource ↗KunstJF/glfBLUP-Procrustes · lines:388-492
Codepublic

ntribute to a greater understanding of high-dimensional data in plant breeding trials. Data and code Scripts to generate the hyperspectral datasets, as well as the results presented in this paper, are available at https://github.com/KunstJF/glfBLUP-Procrustes . The glfBLUP methodology is implemented in an R-package available at https://github.com/KillianMelsen/glfBLUP . The hyperspectral dataset is available upon reasonable request from J. Crossa References Antonio et al. (2022) O. Antonio, M. López, A. Montesinos López, and J. Crossa Multivariate statistical machine learning methods for genomic prediction . Springer , Cham, Switzerland . External Links: ISBN 978-3-030-89009-4 978-3-030-8901

Open resource ↗KillianMelsen/glfBLUP · lines:388-492

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.