ntional (single) scatter correction technique, partial least-squares regression (PLSR) was performed individually pre-processed data. 2. Materials and methods 2.1. Data set The wheat kernel data set used in this study was obtained from the Mendeley repository of open data sets (Wenya, 2016). The data set can also be accessed at https://figshare.com/articles/wheat_kernel_dataset/4252217/1. The data set con- tains NIR spectra and reference protein concentration of 523 wheat kernels. The spectra were measured in the spectral range of 850e1050 nm with a total of 100 wavelengths (nm). In this analysis, the data set was divided into calibration (60%) and test set (40%) using the Kennard-Stone (KS)
Open resource ↗Figshare · 4252217 · pdf-raw-page:2 lines:1-87Unverified paper record
Improved prediction of protein content in wheat kernels with a fusion of scatter correction methods in NIR data modelling
Biosystems engineering. · 1 Mar 2021 · 10.1016/j.biosystemseng.2021.01.003
Abstract
The study aims to test the hypothesis that modelling of near-infrared (NIR) spectroscopic data based on a single scatter correction technique is sub-optimal. Better predictive performance of the multivariate analysis method can be obtained when the information from differently scatter corrected data is jointly used. To demonstrate it, an open-source NIR spectroscopy data set related to protein prediction in wheat kernels was used. Two different pre-processing fusion approaches i.e., sequential and parallel fusion, were used for fusing the complementary information from four different scatter correction techniques, namely standard normal variate (SNV), variable sorting for normalisation (VSN), 2nd derivative, and multiplicative scatter correction (MSC). As a comparison, partial least-squares regression (PLSR) was performed on the SNV pre-processed data. The results showed that fusion of scatter correction can improve the predictive performance of NIR spectroscopic models. The results revealed that both sequential and parallel fusion approaches improved the predictive performance compared to the PLSR performed using a single scatter correction technique. The R²ₚ was improved by up to 3% and the RMSEP was reduced by up to 13% compared to the results obtained with conventional PLSR model developed with a single scatter correction technique.
Plant phenotyping relevance
小麦粒のタンパク質含量という植物形質を対象に、NIRスペクトルの散乱補正融合と予測性能を検証しており、形質取得・推定手法が研究の中心である。
abstractThe study aims to test the hypothesis that modelling of near-infrared (NIR) spectroscopic data based on a single scatter correction technique is sub-optimal.
abstractTwo different pre-processing fusion approaches i.e., sequential and parallel fusion, were used for fusing the complementary information from four different scatter correction techniques
abstractfusion of scatter correction can improve the predictive performance of NIR spectroscopic models.
Code and data availability
The paper's analysis is built entirely on an open NIR spectroscopy dataset of 523 wheat kernels with reference protein content, publicly deposited on Figshare and explicitly linked by the authors. No author analysis code is stated as publicly available.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.