Unverified paper record
A scalable machine learning approach for predicting wheat growth stages with a large national dataset
Field Crops Research. · 1 Feb 2026
Abstract
Accurate wheat growth stage predictions are important for efficient crop management practices, such as when to apply chemical inputs or fertilise. Recent studies have developed accurate machine learning (ML) approaches for predicting the growth stages of wheat and other crops, but these have generally focused on a few key stages and used limited datasets, restricting comprehensive validation across diverse growing seasons and/or regions. This makes it difficult to test their scalability, which is a critical consideration for real-world application. The National Variety Trials (NVT) program represents a key opportunity, providing observations of Zadoks stages since 2005 across the Australian grain belt. To develop a scalable, data-driven approach for predicting wheat growth stages using a national dataset and ML. The dataset contained over 80,000 wheat Zadoks stage observations from the NVT program from 2005 to 2023 across 169 sites in Australia. Models were developed with XGBoost, using 11 weather, remote sensing (RS), genetic and crop management features. Three experiments were designed to evaluate models: 70:30 split, leave-one-year-out (LOYO) and leave-one-site-out (LOSO). The optimal spatial extent was determined by comparing national, regional and subregional models, and a null model was developed to assess quality of predictions if only using features related to temperature. All three spatial extents tested yielded strong results, but the national performed best overall. It showed high accuracy across all three experiments, with strong agreement between observed and predicted Zadoks stages (00−99) (Lin’s concordance correlation coefficient [LCCC] = 0.76–0.80), minimal error (RMSE = 5.9–6.8 stages), and 58–66 % accuracy ±5 Zadoks stages across the three experiments. This model also consistently outperformed the null model, demonstrating that including non-temperature-related features (e.g. solar radiation, variety) led to more accurate growth stage predictions. No other known studies have used such a comprehensive dataset of growth stage observations, both in size and spatiotemporal coverage, to model crop growth stages. The use of this dataset enabled the development and validation of a model that is both accurate and scales well to unseen years and sites. This study therefore highlights the potential for a ML-based operational tool to support crop monitoring across diverse growing seasons and regions. Future work could explore incorporating more RS features and more observations of underrepresented Zadoks stages (e.g. seedling growth).
Plant phenotyping relevance
小麦の生育ステージという植物状態を、機械学習で推定する手法の開発と、年・地点外挿を含む大規模な検証が研究の中心であるため。
abstractTo develop a scalable, data-driven approach for predicting wheat growth stages using a national dataset and ML.
abstractThree experiments were designed to evaluate models: 70:30 split, leave-one-year-out (LOYO) and leave-one-site-out (LOSO).
abstractThe use of this dataset enabled the development and validation of a model that is both accurate and scales well to unseen years and sites.
Code and data availability
公開状態または取得可能な本文経路を確認できませんでした。
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.