← Papers

Unverified paper record

Developing automated machine learning approach for fast and robust crop yield prediction using a fusion of remote sensing, soil, and weather dataset

Environmental Research Communications · 1 Apr 2024 · 10.1088/2515-7620/ad2d02

Abstract

Abstract Estimating smallholder crop yields robustly and timely is crucial for improving agronomic practices, determining yield gaps, guiding investment, and policymaking to ensure food security. However, there is poor estimation of yield for most smallholders due to lack of technology, and field scale data, particularly in Egypt. Automated machine learning (AutoML) can be used to automate the machine learning workflow, including automatic training and optimization of multiple models within a user-specified time frame, but it has less attention so far. Here, we combined extensive field survey yield across wheat cultivated area in Egypt with diverse dataset of remote sensing, soil, and weather to predict field-level wheat yield using 22 Ml models in AutoML. The models showed robust accuracies for yield predictions, recording Willmott degree of agreement, (d > 0.80) with higher accuracy when super learner (stacked ensemble) was used (R 2 = 0.51, d = 0.82). The trained AutoML was deployed to predict yield using remote sensing (RS) vegetative indices (VIs), demonstrating a good correlation with actual yield (R 2 = 0.7). This is very important since it is considered a low-cost tool and could be used to explore early yield predictions. Since climate change has negative impacts on agricultural production and food security with some uncertainties, AutoML was deployed to predict wheat yield under recent climate scenarios from the Coupled Model Intercomparison Project Phase 6 (CMIP6). These scenarios included single downscaled General Circulation Model (GCM) as CanESM5 and two shared socioeconomic pathways (SSPs) as SSP2-4.5and SSP5-8.5during the mid-term period (2050). The stacked ensemble model displayed declines in yield of 21% and 5% under SSP5-8.5 and SSP2-4.5 respectively during mid-century, with higher uncertainty under the highest emission scenario (SSP5-8.5). The developed approach could be used as a rapid, accurate and low-cost method to predict yield for stakeholder farms all over the world where ground data is scarce.

Plant phenotyping relevance

圃場レベルの小麦収量という植物形質を、リモートセンシング等とAutoMLで推定する手法を開発・検証しており、収量取得・予測ワークフローが中心である。

titleDeveloping automated machine learning approach for fast and robust crop yield prediction using a fusion of remote sensing, soil, and weather dataset
abstractHere, we combined extensive field survey yield across wheat cultivated area in Egypt with diverse dataset of remote sensing, soil, and weather to predict field-level wheat yield using 22 Ml models in AutoML.
abstractThe trained AutoML was deployed to predict yield using remote sensing (RS) vegetative indices (VIs), demonstrating a good correlation with actual yield (R 2 = 0.7).

Code and data availability

The paper's authors developed an H2O AutoML workflow in R for wheat yield prediction and explicitly state the full script is publicly hosted on GitHub. This is a paper-specific, publicly available analysis code asset. The data availability statement only says data are within the article/supplements, so no separate phen

Codepublic

web GUI, H2O AutoML is also accessible in Python, R, Java, and Scala. The technique is entirely automated, but many of the settings are made available to the user as parameters so that some parts of the modelling phases can be changed. In our case, we developed H2OAutoML in R language and the full script is hosted on GitHub at https://github.com/DrAhmedKheir/H2O_ AutoML.git. 2.3.1. Dataset preprocessing and AutoML training Currently, all H2O supervised learning algorithms offer the same kind of automatic data-preprocessing as H2O AutoML. Categorical data can be handled natively because H2O tree-based models (Gradient Boosting Machines, Random Forests) provide group-splits on categorical

Open resource ↗DrAhmedKheir/H2O_ · pdf-layout-page:6 lines:1-35

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.