← Papers

Unverified paper record

High-throughput Cotton Phenotyping Big Data Pipeline Lambda Architecture Computer Vision Deep Neural Networks

arXiv (Cornell University) · 9 May 2023 · 10.48550/arxiv.2305.05423

Abstract

In this study, we propose a big data pipeline for cotton bloom detection using a Lambda architecture, which enables real-time and batch processing of data. Our proposed approach leverages Azure resources such as Data Factory, Event Grids, Rest APIs, and Databricks. This work is the first to develop and demonstrate the implementation of such a pipeline for plant phenotyping through Azure's cloud computing service. The proposed pipeline consists of data preprocessing, object detection using a YOLOv5 neural network model trained through Azure AutoML, and visualization of object detection bounding boxes on output images. The trained model achieves a mean Average Precision (mAP) score of 0.96, demonstrating its high performance for cotton bloom classification. We evaluate our Lambda architecture pipeline using 9000 images yielding an optimized runtime of 34 minutes. The results illustrate the scalability of the proposed pipeline as a solution for deep learning object detection, with the potential for further expansion through additional Azure processing cores. This work advances the scientific research field by providing a new method for cotton bloom detection on a large dataset and demonstrates the potential of utilizing cloud computing resources, specifically Azure, for efficient and accurate big data processing in precision agriculture.

Plant phenotyping relevance

綿花の花の検出という植物器官の表現型取得を対象に、クラウド型ビッグデータ処理パイプラインとYOLOv5による画像解析手法を開発・評価しており、方法が研究の中心である。

abstractThis work is the first to develop and demonstrate the implementation of such a pipeline for plant phenotyping through Azure's cloud computing service.
abstractThe proposed pipeline consists of data preprocessing, object detection using a YOLOv5 neural network model trained through Azure AutoML, and visualization of object detection bounding boxes on output images.

Code and data availability

The paper describes a cotton field dataset (9,018 sliced images) and a YOLOv5 model trained via Azure AutoML, but no block contains a public deposit, availability statement, or authors' URL for the dataset, images, code, or model checkpoints. The only URLs present are generic library documentation and citations, which,

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.