Unverified paper record
Rapid Determination of Soybean Protein Content by Near-Infrared Spectroscopy Coupled with Multi-Learner Ensemble Wavelength Selection.
Foods (Basel, Switzerland) · 15 May 2026 · 10.3390/foods15101755
Abstract
Soybean protein content is a key indicator of nutritional value and quality grade, and its determination is important for quality evaluation and cultivar selection. To overcome the time-consuming and costly limitations of conventional chemical assays, this study proposed a multiple linear learner ensemble importance-score wavelength selection (MLLEISWS) method to identify informative wavelengths from soybean near-infrared spectra and establish a partial least squares (PLS) model. MLLEISWS was compared with competitive adaptive reweighted sampling, successive projections algorithm, and uninformative variable elimination. Shapley additive exPlanations (SHAP) were applied to the MLLEISWS algorithm to interpret the selected wavelengths. Results showed that the PLS model developed using MLLEISWS achieved the best performance. With only 29 selected wavelengths, the coefficients of determination for the training and test sets reached 0.941 and 0.933, respectively. Root mean square errors were 0.490% and 0.514%, relative root mean square errors were 1.32% and 1.37%, and residual predictive deviation was 3.863, indicating predictive accuracy and stability. SHAP analysis showed that the selected wavelengths were located in protein-related spectral regions and corresponded to overtone and combination bands information from functional groups. MLLEISWS effectively reduced variable dimensionality while maintaining model performance.
Plant phenotyping relevance
大豆種子のタンパク質含量という植物形質を対象に、近赤外分光法と波長選択・PLSモデルを開発、比較評価しており、形質取得・推定手法が研究の中心である。
abstractthis study proposed a multiple linear learner ensemble importance-score wavelength selection (MLLEISWS) method to identify informative wavelengths from soybean near-infrared spectra and establish a partial least squares (PLS) model.
abstractMLLEISWS was compared with competitive adaptive reweighted sampling, successive projections algorithm, and uninformative variable elimination.
Code and data availability
The article describes soybean NIR spectra and protein measurements analyzed in MATLAB, but no public dataset, code, or model deposit is mentioned. The Data Availability Statement only offers inquiries via the corresponding author, and no authors' public URL is provided.
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.