Unverified paper record
Image‐based and biochemical multimodal phenotyping for explainable classification of chia ( Salvia hispanica L.) genotypes
Journal of the Science of Food and Agriculture · 31 Aug 2026 · 10.1002/jsfa.71038
Abstract
Abstract BACKGROUND This study developed an explainable machine learning framework integrating morphological, color, and biochemical characteristics for classifying chia ( Salvia hispanica L.) genotypes. A dataset was assembled from 1200 seed images spanning four genotypes, from which 17 morphological and color features were extracted. These were complemented by six sample‐level biochemical traits – crude protein, fat, ash, fiber, carbohydrate, and total sugar – obtained from the corresponding experimental‐unit seed sample, resulting in a total of 23 variables in the integrated dataset. The dataset was evaluated comparatively with 10 machine learning algorithms under repeated 10‐fold cross‐validation, with all preprocessing confined to each training fold to avoid data leakage. RESULTS The highest performance was obtained with XGBoost, reaching 86.99% accuracy, a Matthews correlation coefficient of 0.820, a receiver operating characteristic (ROC) area of 0.975, and a precision–recall curve (PRC) area of 0.933; Simple Logistic followed closely at 86.85% accuracy, with comparable ROC and PRC areas (0.974 and 0.933). Significant differences among the algorithms were confirmed by the Friedman test ( P = 2.47 × 10 −120 ), with post hoc comparisons placing XGBoost and Simple Logistic within the same top‐performing group. Protein, fiber, ash, and fat were the most influential biochemical traits, while hue and saturation among color parameters and shape index and geometric mean diameter among morphological features also contributed appreciably. The G1 genotype, which showed comparatively high protein (27.62%) and fiber (40.62%) contents, was the most consistently distinguished class, with XGBoost and Simple Logistic achieving F‐measures of 0.954 and 0.955, respectively, whereas greater phenotypic overlap between G2 and G3 resulted in more frequent mutual misclassifications. CONCLUSION These findings indicate that multimodal phenotyping, coupled with explainable machine learning, offers a practical and biologically interpretable decision‐support approach for chia genotype classification. © 2026 The Author(s). Journal of the Science of Food and Agriculture published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.
Plant phenotyping relevance
画像から種子の形態・色形質を抽出し、生化学形質と統合したマルチモーダル表現型解析・機械学習分類法の開発と比較評価が研究の中心であるため。
abstractThis study developed an explainable machine learning framework integrating morphological, color, and biochemical characteristics for classifying chia ( Salvia hispanica L.) genotypes.
abstractfrom which 17 morphological and color features were extracted.
Code and data availability
The paper describes chia seed image scans, extracted morphological/color features, biochemical measurements, and ML analysis, but no public dataset, image repository, or author code URL is provided. The data availability statement only offers data from the corresponding author upon reasonable request; supporting信息 is a
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.