Unverified paper record
Rapid Analysis of Caffeine, Protein and Trigonelline in Ugandan Arabica Coffee Using NIRS and Machine Learning Algorithms.
Plants (Basel, Switzerland) · 9 Jul 2026 · 10.3390/plants15142117
Abstract
Coffee is a major export earner for Uganda, raking in over USD 2 billion in 2025. The global price of coffee is tagged to the perceived quality in the cup which in turn is affected by the chemical composition of the green bean. Breeding for market-preferred Arabica coffee varieties is a major objective of coffee breeding programs. Determination of coffee bean chemical constituents is routinely done through expensive, slow and tedious laboratory procedures, making it unsustainable of resource-limited public sector coffee breeding programs. Here, we demonstrate the use of near-infrared spectroscopy (NIRS) and the machine learning algorithms partial least squares (PLS), random forest (RF) and support vector machine (SVM) for the prediction of caffeine, protein and trigonelline in Arabica coffee. NIRS provides a fast, accurate and reliable method of simultaneously predicting multiple sample constituents. Ripe coffee cherries were picked from 172 farmers' fields, air dried in the laboratory at room temperature and processed to green beans. NIRS spectra were taken on the milled green bean at 400-2500 nm, with a 0.5 nanometer (nm) step. Reference data for caffeine, protein and trigonelline were collected on the same sample scanned with NIRS. A set of 12 spectral pretreatments were applied prior to making calibrations with the PLS, RF and SVM algorithms and 70% of the data as a training set and 30% as a test set. Caffeine content of reference samples ranged from 1.94-3.0 g/100 g, protein content ranged from 11.16-15.94% while trigonelline ranged from 0.94-1.23 g/100 g. The best calibrations for all algorithms and analytes were obtained using raw (untreated) spectra, which gave the same results as the Savitzky-Golay (SG) pretreatment. For caffeine, the best model (R 2 p = 0.89, RMSEP = 0.007, RPD = 3.34) was obtained with the SVM algorithm, while for protein, the best model (R 2 p = 0.98, RMSEP = 0.14, RPD = 6.92) was obtained using the PLS algorithm. Finally, for trigonelline, all three models had very high prediction accuracies (R 2 p = 0.98-0.99, RMSEP = 0.007-0.009, RPD = 8.53-10.52). Collectively, these results demonstrate the potential of using NIRS for rapid and simultaneous prediction of coffee green bean constituents to aid selection decisions.
Plant phenotyping relevance
コーヒー生豆の化学的形質を対象に、NIRSと機械学習による予測モデルを開発・検証しており、形質取得・推定法が研究の中心である。育種選抜への利用も明示されている。
abstractwe demonstrate the use of near-infrared spectroscopy (NIRS) and the machine learning algorithms partial least squares (PLS), random forest (RF) and support vector machine (SVM) for the prediction of caffeine, protein and trigonelline in Arabica coffee.
abstractCollectively, these results demonstrate the potential of using NIRS for rapid and simultaneous prediction of coffee green bean constituents to aid selection decisions.
Code and data availability
The paper's NIRS spectra (172 samples, 400–2500 nm) and HPLC reference measurements for caffeine, protein and trigonelline are the core phenotyping data, but the Data Availability Statement says they are only available from the authors on request due to institutional restrictions. The MDPI supplementary file (allowed,
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.