Unverified paper record
Explainable TabPFN-Based Machine Learning for Single-Plant Yield Estimation and Trait Prioritization in Faba Bean (Vicia faba L.)
Agronomy · 28 Aug 2026 · 10.3390/agronomy16171653
Abstract
Faba bean yield reflects complex relationships among genotype, environment, and agronomic traits. This study evaluated an explainable Tabular Prior-data Fitted Network (TabPFN) framework for estimating plot-mean single-plant yield and prioritizing traits using 398 plot-level observations, 13 measured agronomic predictors, and six derived features. On the reference 80/20 split, TabPFN achieved the best values for all four test metrics (R2 = 0.8746, RMSE = 1.9132 g plant−1, MAE = 1.0819 g plant−1, and MAPE = 8.16%). The Friedman test detected differences among the six models (χ2(5) = 16.75, p = 0.005); Nemenyi comparisons distinguished TabPFN from HistGradientBoosting and SVR, whereas the Holm-corrected Wilcoxon analysis confirmed only the TabPFN–SVR difference. Across 10 repeated 80/20 splits, TabPFN obtained the highest mean test R2 (0.8614 ± 0.0691), ranked first in eight splits, and produced a higher R2 than every tuned baseline in at least eight splits. SHAP, permutation importance, and LOCO analyses emphasized pod-, seed-, and biomass-related predictors. Repeated-split ablation showed that derived features improved TabPFN consistently, whereas removing selected target-proximal yield variables reduced performance for every model. The framework is therefore a harvest-time trait-estimation and trait-prioritization tool rather than an early-season forecasting system. Notably, TabPFN achieved this performance without the 100-trial Optuna search used for each baseline; only n_estimators was screened over four prespecified values.
Plant phenotyping relevance
単一個体収量を推定し、形質優先順位付けを行う機械学習フレームワークを評価・比較しており、植物形質抽出手法が研究の中心である。
abstractThis study evaluated an explainable Tabular Prior-data Fitted Network (TabPFN) framework for estimating plot-mean single-plant yield and prioritizing traits
abstractThe framework is therefore a harvest-time trait-estimation and trait-prioritization tool rather than an early-season forecasting system.
Code and data availability
The supplied blocks describe a 398-plot faba bean agronomic dataset and TabPFN/SHAP analysis, but contain no public data deposit, repository, or code availability statement. The only data source mentioned is gene bank seed accessions, not a paper-specific public dataset or code URL.
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.