← Papers

Unverified paper record

Cyberinfrastructure for machine learning applications in agriculture: experiences, analysis, and vision.

Frontiers in artificial intelligence · 23 Jan 2025 · 10.3389/frai.2024.1496066

Abstract

Introduction Advancements in machine learning (ML) algorithms that make predictions from data without being explicitly programmed and the increased computational speeds of graphics processing units (GPUs) over the last decade have led to remarkable progress in the capabilities of ML. In many fields, including agriculture, this progress has outpaced the availability of sufficiently diverse and high-quality datasets, which now serve as a limiting factor. While many agricultural use cases appear feasible with current compute resources and ML algorithms, the lack of reusable hardware and software components, referred to as cyberinfrastructure (CI), for collecting, transmitting, cleaning, labeling, and training datasets is a major hindrance toward developing solutions to address agricultural use cases. This study focuses on addressing these challenges by exploring the collection, processing, and training of ML models using a multimodal dataset and providing a vision for agriculture-focused CI to accelerate innovation in the field. Methods Data were collected during the 2023 growing season from three agricultural research locations across Ohio. The dataset includes 1 terabyte (TB) of multimodal data, comprising Unmanned Aerial System (UAS) imagery (RGB and multispectral), as well as soil and weather sensor data. The two primary crops studied were corn and soybean, which are the state's most widely cultivated crops. The data collected and processed from this study were used to train ML models to make predictions of crop growth stage, soil moisture, and final yield. Results The exercise of processing this dataset resulted in four CI components that can be used to provide higher accuracy predictions in the agricultural domain. These components included (1) a UAS imagery pipeline that reduced processing time and improved image quality over standard methods, (2) a tabular data pipeline that aggregated data from multiple sources and temporal resolutions and aligned it with a common temporal resolution, (3) an approach to adapting the model architecture for a vision transformer (ViT) that incorporates agricultural domain expertise, and (4) a data visualization prototype that was used to identify outliers and improve trust in the data. Discussion Further work will be aimed at maturing the CI components and implementing them on high performance computing (HPC). There are open questions as to how CI components like these can best be leveraged to serve the needs of the agricultural community to accelerate the development of ML applications in agriculture.

Plant phenotyping relevance

農業向けサイバーインフラの開発が中心で、UAS画像処理パイプラインとMLモデルにより作物の生育ステージおよび収量を推定しており、植物形質の取得・抽出方法が実質的に扱われている。

abstractThis study focuses on addressing these challenges by exploring the collection, processing, and training of ML models using a multimodal dataset and providing a vision for agriculture-focused CI to accelerate innovation in the field.
abstractThe data collected and processed from this study were used to train ML models to make predictions of crop growth stage, soil moisture, and final yield.
abstractThese components included (1) a UAS imagery pipeline that reduced processing time and improved image quality over standard methods

Code and data availability

The paper describes a 1 TB multimodal UAS/sensor dataset and CI pipelines, but no public deposit, availability statement, or authors' URL for the dataset, code, or models is provided in the supplied blocks. The only URLs are the license and cited external references (MLCommons Croissant announcement, Purdue extension),

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.