← Papers

Unverified paper record

OPTIMIZING PRE-PROCESSING OF NEAR INFRARED SPECTRA FOR PHENOMIC PREDICTION USING SINGULAR VALUE DECOMPOSITION

13 May 2026 · 10.64898/2026.05.10.724118

Abstract

Phenomic prediction (PP) is a genetic value prediction method based on near infrared spectroscopy (NIRS). Spectra pre-processing is a key step in the analysis pipeline of PP and generally involves chemometrics methods. However, the choice of pre-processing is usually done either arbitrarily or through a search of the optimal set of methods and associated parameters. In this study, we propose to implement a singular value decomposition (SVD) step in the pre-processing pipeline where genetic values of spectra are estimated on a set of principal components instead of individual wavelengths. This way, estimations are based on a few informative, orthogonal and interpretable features of spectra instead of many correlated, uninformative wavelengths. We tested this pre-processing method on five datasets representing four plant species (maize, rice, sorghum and grapevine). Results show that estimating genetic values on components of raw spectra, that are not weighted by their eigenvalues, performs as well as doing it on spectra pre-processed with the best classical chemometrics methods in most cases, while requiring less parameter optimization. Moreover, this SVD step opens up possibilities for better understanding and selecting parts of the spectral information that are relevant for PP. Plain language summary Cultivated plants are the result of a breeding process during which their genetic values are used to select those to breed. Estimating these values requires heavy experimental means and is time consuming. Phenomic prediction is a low cost and high throughput method that is increasingly being used for this purpose. It often uses, as predictors, near infrared spectroscopy measurements that are easy to collect and thus routinely used in many species. However, near infrared spectra generally require pre-processing before being used in prediction. Currently used pre-processing methods arise from the chemometrics community, and still deserve a better in-depth appropriation by geneticists. In this study, we propose a pre-processing approach that performs as well as the best chemometrics pre-processing generally used, reduces computation time, and allows for a better understanding of what parts of spectral information are relevant for prediction. Core Ideas The SVD-based pre-processing performs as well as the best performing classical chemometrics pre-processing in most cases Using the SVD-based pre-processing reduces computing time of genetic value estimation and requires less parameter optimization than using classical chemometrics pre-processing Spectra are composed of chemical and physical information and classical pre-processing methods remove the physical part of the signal It is likely that chemical information is the most important for phenomic prediction even though physical information remains valuable Performance of the SVD-based pre-processing is likely due to a good estimation of the genetic part of spectra and the conservation of physical information of spectra

Plant phenotyping relevance

植物のNIRSスペクトルから遺伝的価値を推定するフェノミック予測について、SVDベースの前処理法を提案し、複数植物種のデータセットで既存法と比較検証しているため、フェノタイピング手法が中心である。

abstractIn this study, we propose to implement a singular value decomposition (SVD) step in the pre-processing pipeline where genetic values of spectra are estimated on a set of principal components instead of individual wavelengths.
abstractWe tested this pre-processing method on five datasets representing four plant species (maize, rice, sorghum and grapevine).
abstractResults show that estimating genetic values on components of raw spectra, that are not weighted by their eigenvalues, performs as well as doing it on spectra pre-processed with the best classical chemometrics methods in most cases

Code and data availability

The supplied blocks describe five NIRS phenomic-prediction datasets and SVD pre-processing methods, but contain no data availability statement, no public deposit of spectra/phenotypes, and no author code or repository URL. Only generic R packages (lme4, sommer, ggplot2, tidyverse) are cited, which are not paper-qualify

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.