← Papers

Unverified paper record

Optimizing chlorophyll content prediction in tea leaves via spectral transformations and deep learning.

BMC Plant Biology · 1 Dec 2025 · 10.1186/s12870-025-07863-2

Abstract

Accurate estimation of chlorophyll contents from spectral reflectance is necessary for monitoring plant physiological status and for supporting precision agriculture. This study, which uses four machine learning models (1D Convolutional Neural Network (1D-CNN), Self-Supervised Learning (SSL), Vision Transformer (ViT), and Conformer), elucidates the effects of four preprocessing techniques on the performance of chlorophyll content prediction: Original Reflectance (OR), Continuum Removal (CR), De-trending (DT), and Standard Normal Variate (SNV). Reflectance data were collected from tea leaves (Camellia sinensis) and were analysed using ten-fold cross-validation. Correlation analysis revealed that SNV and DT enhanced the spectral sensitivity to chlorophyll content, particularly around the chlorophyll absorption regions (450-500 nm and 650-700 nm), whereas CR emphasized negative correlation in the visible spectrum. Prediction results demonstrated that the SSL model combined with SNV preprocessing achieved the highest accuracy (R² = 0.82, RPD = 2.37), outperforming other model-preprocessing combinations. The 1D-CNN model performed best with DT, leveraging local spectral features, whereas ViT and Conformer models benefited most from CR, which emphasizes absorption depth and spectral shape. These results highlight that the optimal preprocessing method depends on the model architecture, and that proper pairing between preprocessing and modelling approaches is crucially important for maximizing prediction performance. The study results underscore the importance of customized preprocessing strategies for hyperspectral analysis and provide practical insights for improving biochemical trait estimation in plant phenotyping.

Plant phenotyping relevance

茶葉のスペクトル反射からクロロフィル含量を推定する植物表現型測定法について、前処理と複数の深層学習モデルを比較・検証しており、方法論が研究の中心です。

abstractAccurate estimation of chlorophyll contents from spectral reflectance is necessary for monitoring plant physiological status
abstractThis study, which uses four machine learning models (1D Convolutional Neural Network (1D-CNN), Self-Supervised Learning (SSL), Vision Transformer (ViT), and Conformer), elucidates the effects of four preprocessing techniques on the performance of chlorophyll content prediction
abstractPrediction results demonstrated that the SSL model combined with SNV preprocessing achieved the highest accuracy (R² = 0.82, RPD = 2.37)

Code and data availability

The paper's hyperspectral reflectance measurements and chlorophyll phenotype data (3,120 tea leaf samples) are not publicly deposited; the Data Availability Statement says they are available only by contacting the corresponding author. No author code, models, or public repository URLs are provided.

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.