← PhenoCode Atlas

Unverified paper discovery

Plant phenotyping methods.

植物形質を測っただけの研究ではなく、フェノタイピング手法の開発・検証・実質的利用・ベンチマーク・方法レビューとの関連性が見つかった論文を中心に表示します。

表示条件: dataset_or_benchmark条件を解除 ×
515 papers · 上位300件を表示 · code / dataset availability confirmedLatest completed run · 2016-01-01 – 2026-09-13

自動判定された未検証候補です。Catalogへの掲載にはキュレーター承認が必要です。

Code / dataset availability confirmedOpenAlex · checked 15 Sept 2026
Published7 Sept 2026Plant PhenomicsCited by 0 · OpenAlex ↗

FG-LCNet: A two-stage foreground-guided network for whole-tree litchi counting

Field / plotFruitWhole plant / canopy / plot / fieldCountingObject detectionFruit / seed / panicle traits

Accurate litchi counting from whole-tree images is essential for yield estimation, orchard management, and plant phenotyping, but remains challenging in real orchards because fruits occur in dense, heavily occluded clusters and vary markedly in scale, illumination, and appearance across ripening stages, particularly when green fruits resemble surrounding foliage. Existing methods have shown promise, but their robustness in complex orchard environments remains limited. To address these challenges, we propose FG-LCNet, a two-stage foreground-guided litchi counting framework. In the first stage, an enhanced fruit-cluster detector improves the localization of small and ambiguous clusters under complex canopy backgrounds. In the second stage, the detected foreground regions are fed into a density-regression network with hybrid attention, while a consistency-based training strategy is introduced to improve robustness to appearance and illumination variations. To support this study, a large-scale litchi counting dataset was established, consisting of 1,126 whole-tree images collected from five orchards and spanning three ripening stages, with approximately 120,000 fruit-level dot annotations and more than 20,000 cluster-level bounding boxes. FG-LCNet achieved the best overall counting performance, with an MAE of 7.44 and an RMSE of 11.01. It showed clear advantages in high-density fruit-cluster scenarios and cross-orchard validation, while maintaining competitive results across orchard-region and maturity-stage subsets. The framework further retained inference efficiency suitable for practical deployment. These results indicate that FG-LCNet provides an effective solution for robust litchi counting and offers potential for other clustered fruit-counting tasks.

Why it matches plant phenotyping methods果実数という植物器官形質を whole-tree 画像から推定する二段階画像解析手法を開発し、データセット構築と交差果樹園検証まで行っており、表現型取得・抽出法が中心である。

abstractwe propose FG-LCNet, a two-stage foreground-guided litchi counting framework.
Reproduction assets foundThe paper's implementation code is explicitly stated as publicly available at the authors' GitHub repository (FG-LCNet). The litchi counting dataset (1,126 whole-tree images with ~120,000 dot annotations and 20,000+ bounding boxes) is not yet fully public: a ~100-image annotated subset is promised upon acceptance, and,
Code · publicdustry Technology Research System (CARS-32-21), Hainan Modern Agricul- 655 tural Industry Technology System (HNARS-08-G02). 656 Conflicts of Interest 657 The authors declare that there is no conflict of interest regarding the publication of this article. 658 Data Availability 659 The implementation code is publicly available at https://github.com/johnhamtom/FG-LCNet . 660 Upon acceptance, a representative subset of approximately 100 annotated litchi images will be 661 released to support reproducibility and preliminary benchmarking. The full dataset is being further 662 organized for future release. Before full release, the complete dataset can be obtained from the 663 corresponding authorOpen asset ↗johnhamtom/FG-LCNetpdf-raw-page:28 lines:1-81
Code / dataset availability confirmedEurope PMC · Crossref · checked 15 Sept 2026
Published4 Sept 2026Springer Science and Business Media LLC

Deep Learning-Based Crop Disease Detection Using EfficientNet-B3 for Smart Agriculture

CottonField / plotRGB / grayscaleLeafClassificationDisease symptoms / severity

Abstract Plant diseases substantially reduce global crop yields, and cotton production is particularly vulnerable to field-acquired variability in symptom appearance, background clutter, and illumination changes that limit the reliability and scalability of expert visual inspection. This study aimed to develop an accurate, computationally efficient, and explainable framework for real-time cotton leaf disease recognition that is suitable for deployment on resource-constrained edge devices. Using the SAR-CLD-2024 dataset (322 RGB images captured under natural agricultural conditions across seven categories, including healthy and diseased leaves), images were preprocessed via resizing and normalization and augmented online in the training set (random rotations, flips, brightness/contrast adjustments, and random cropping). An EfficientNet-B3 backbone initialized with ImageNet-pretrained weights was fine-tuned using categorical cross-entropy loss and Adam optimization, with early stopping, checkpointing, regularization, and a fixed-seed 70/15/15 train–validation–test partition to enhance reproducibility and reduce leakage. Performance was evaluated on an independent test set using accuracy, precision, recall, F1-score, MCC, balanced accuracy, Cohen’s kappa, confusion matrix, multi-class ROC/AUC, and precision–recall analysis, alongside computational benchmarking (parameters, FLOPs, memory, and inference latency) and comparative experiments against contemporary CNN, lightweight, and transformer-based models. The model showed stable convergence over 30 epochs with a small training–validation gap, predominantly correct predictions with limited confusion among visually similar classes, consistently high precision–recall behavior under moderate class imbalance, and stable performance across repeated runs with low variability and a tight confidence interval. Grad-CAM heatmaps localized necrotic lesions, discoloration, and infected tissues while largely ignoring background, and failure cases were associated with early-stage symptoms, occlusion, shadows, and inter-class similarity. Overall, the framework provides a reproducible, interpretable, and efficient solution for cotton leaf disease classification with practical implications for trustworthy, low-latency, on-device decision support in precision agriculture.

Why it matches plant phenotyping methods綿葉の病徴を画像から分類する深層学習手法の開発が中心で、独立テスト、比較評価、計算性能評価、Grad-CAMによる病徴局在化を実施しているため、植物病害フェノタイピング手法に該当する。

abstractThis study aimed to develop an accurate, computationally efficient, and explainable framework for real-time cotton leaf disease recognition
Reproduction assets foundThe paper's Data Availability statement explicitly names the SAR-CLD-2024 cotton leaf dataset used for all experiments as publicly available on Kaggle with a direct URL. No author analysis code or trained model checkpoint is deposited.
Dataset · publicntribute to the development of fully automated, scalable, and real-time smart agriculture systems. Declaration Funding Datta Meghe Institute of Higher Education and Research Wardha, Maharashtra, India Data Availability: The SAR-CLD-2024 cotton leaf dataset used in this study is publicly available through the Kaggle platform at: https://www.kaggle.com/datasets/pantho12/sar-cld-2024-dataset-for-cotton This dataset includes annotated images of various cotton leaf diseases collected under diverse environmental conditions. All data utilized in this work are freely accessible, and the data processing methodology has been described in detail to facilitate reproducibility. Conflict of interest The aOpen asset ↗Kaggle · SAR-CLD-2024pdf-raw-page:32 lines:1-38
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published2 Sept 2026

A methodological framework for the standardised evaluation of olive genetic resources: GEN4OLIVE harmonized protocols

OliveField / plotFruitWhole plant / canopy / plot / fieldMorphology / geometry measurementStress / disease detectionYield / biomass estimationGrowth / development / phenologyStress response / toleranceYield / yield components

Background Olive ( Olea europaea L.) breeding initiatives rely heavily on the extensive and correct characterisation of genetic resources to successfully achieve their goals, such as addressing climate change and emerging diseases challenges. However, the historical lack of standardised phenotyping protocols across multi-environment trials has severely hindered data interoperability and large-scale comparative analyses. Methods Within the Horizon 2020 GEN4OLIVE project, five international olive germplasm banks established a consensus-based methodological framework to systematically evaluate over 500 olive cultivars. We harmonised 14 evaluation protocols covering six fundamental dimensions: phenological and agronomic traits, abiotic stress resilience, biotic stress resilience, olive oil yield and chemical quality, table olive quality assessment, and morphological characterisation and photography. While most protocols were adapted from previously published literature to ensure ease of implementation across different facilities, novel methodologies for frost tolerance and standardised photography were developed de novo. Results The implementation of these consensus methods across five countries proved highly successful. This methodological framework enabled the generation of the largest harmonised, publicly available dataset on olive genetic resources to date, effectively making possible the correct comparation and ranking of the olive cultivars based on their specific characteristics. Conclusions This compendium of methods provides a robust, highly replicable reference point for the standardisation of olive germplasm characterisation and use of shared benchmark cultivars as an effective way for data normalization and comparation. It facilitates future global pre-breeding efforts, ensures international data interoperability, and supports the discovery of resilient cultivars to secure the future of the olive sector.

Why it matches plant phenotyping methodsオリーブ遺伝資源の標準化フェノタイピングプロトコルを体系化し、複数機関で実装・検証して大規模データセットを生成した方法論中心の研究である。

abstractthe historical lack of standardised phenotyping protocols across multi-environment trials has severely hindered data interoperability and large-scale comparative analyses.
Reproduction assets foundThe article declares two paper-specific public assets: the GEN4OLIVE phenotypic dataset from evaluating over 500 olive accessions across five germplasm banks, hosted on the project's Olive Varieties Database, and a Zenodo-deposited methodological handbook (Extended Data) containing the 14 protocols, visual assessment,
Dataset · publicData and software availability The phenotypic dataset generated from the evaluation of over 500 olive varieties across the five Mediterranean germplasm banks using this compendium of protocols and methodologies, is publicly available via the GEN4OLIVE project repository. • Repository: GEN4OLIVE Olive Varieties Database. • Link: https://www.uco.es/ucolivo/gen4olive/olivevarieties (GEN4OLIVE Database, 2025). Page 8 of 15 Open Research Europe 2026, 6:322 Last updated: 14 SEP 2026Open asset ↗GEN4OLIVE Olive Varieties Databasepdf-raw-page:8 lines:1-44
Code / dataset availability confirmedOpenAlex · arXiv · checked 5 Sept 2026
Published28 Aug 2026arXiv (Cornell University)Cited by 0 · OpenAlex ↗

Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait Extraction

MaizeTomatoLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurementCalibration / preprocessing2D/3D reconstructionGrowth / time-series analysisArchitecture / morphology / geometryPlant / canopy height

High-throughput phenotyping depends on accurate 3D reconstruction of plants across growth stages, yet the development and evaluation of temporal completion methods are limited by the lack of datasets with complete geometric ground truth. To address this challenge, we introduce SynthCrop4D, a procedurally generated synthetic dataset of temporally evolving plant point clouds that provides controllable noise, occlusion, and complete plant geometry for benchmarking reconstruction methods. Using this dataset, we evaluate a two-stage pipeline that combines spatial denoising and temporal point cloud completion. First, a denoising module removes structural artifacts from raw laser-scanned point clouds. The resulting data are then processed by an Adaptive Temporal PoinTr model that reconstructs the current growth stage (t) using information from the previous stage (t-1), enabling recovery of regions missing due to self-occlusion. We evaluate the proposed framework on both SynthCrop4D and the real-world Pheno4D dataset (tomato and maize) under settings with and without denoising. Results show that denoising substantially improves reconstruction quality, with the best configuration achieving a Chamfer Distance of 0.0061 on SynthCrop4D (Temporal PoinTr + Mamba-DG) and an F-Score of 0.2080 on Pheno4D (Vanilla PoinTr + Mamba-DG). We further demonstrate the use of completed point clouds for phenotypic trait extraction, including plant height, canopy width, and convex hull volume, obtaining hull-volume MAEs of 0.021 on synthetic data and 0.343 on real data. Together, SynthCrop4D and the proposed pipeline provide a benchmark and methodology for temporal plant reconstruction and high-throughput crop phenotyping.

Why it matches plant phenotyping methods植物の3D点群再構成・時系列補完を開発し、合成データセットと実データで性能検証するとともに、草丈・群落幅・凸包体積を抽出する手法を中心に扱っている。

abstractwe introduce SynthCrop4D, a procedurally generated synthetic dataset of temporally evolving plant point clouds that provides controllable noise, occlusion, and complete plant geometry for benchmarking reconstruction methods.
Reproduction assets foundThe paper's authors explicitly state that source code, implementation details, and pre-trained model weights are publicly available on GitHub, and that the paper-specific SynthCrop4D synthetic dataset can be reproduced via scripts in that codebase. Pheno4D is a cited prior public dataset, not a paper-specific asset.
Code · publicThe source code, implementation details, and pre-trained model weights for this study are publicly available on GitHub at https://github.com/Mrudul2006/3d_plant-reconstruction .Open asset ↗Mrudul2006/3d_plant-reconstructionlines:1724-1761
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published27 Aug 2026Cited by 0 · OpenAlex ↗

An annotated dataset of soybean root nodules for deep learning-based object detection

SoybeanRootObject detection

Abstract Technological advances have expanded the adoption of digital technologies in agriculture, helping to reduce labour effort, increase profitability, improve crop efficiency and productivity, enhance product quality, mitigate environmental impacts, and promote human health. This context also extends to soybean farming, a sector of major economic importance in Brazil. Most importantly, Brazil has the global leadership in soybean production with biological nitrogen fixation (BNF) replacing chemical fertilisers. The research and evaluation of BNF is limited by manual counting of nodules, a time-consuming procedure. This study presents SoyNodules, designed for the automatic identification of soybean nodules, consisting of a dataset of images. The dataset includes 1,701 images acquired under controlled conditions: 1,662 images of soybean roots with nodules and 39 images of isolated nodules without roots. A total of 49,210 nodule instances are manually annotated with bounding boxes. SoyNodules was designed to promote reuse and interoperability in alignment with the FAIR principles (Findable, Accessible, Interoperable, Reusable) and to support the development, training, and evaluation of computer vision and deep learning methods for precision agriculture.

Why it matches plant phenotyping methods大豆根粒を自動識別する画像データセットであり、手作業計数の代替となる植物器官形質の抽出・評価を支援する方法論的データセット。

abstractThis study presents SoyNodules, designed for the automatic identification of soybean nodules, consisting of a dataset of images.
Reproduction assets foundThe paper is a data descriptor for SoyNodules, an annotated dataset of 1,701 soybean root/nodule images with 49,210 bounding-box annotations, publicly deposited on Zenodo with a DOI. The same repository also hosts the authors' annotation-format conversion script (AnyLabeling to Pascal VOC/COCO), per the Code Availabil­
Dataset · publicThe SoyNodules dataset, released as version 1.0, is publicly available on Zenodo [28] at https://doi.org/10.5281/zenodo.22081914.Open asset ↗Zenodo · 10.5281/zenodo.22081914pdf-page:9 lines:1-43
Code / dataset availability confirmedCrossref · checked 11 Sept 2026
Published25 Aug 2026Earth System Science DataCited by 0 · OpenAlex ↗

NortheastChinaMaizeYield10m: a 10 m resolution maize yield dataset for Northeast China (2019–2024) generated via a mechanistically interpretable, field-label-free framework

MaizeField / plotWhole plant / canopy / plot / fieldGrowth / time-series analysisYield / biomass estimationYield / yield components

Abstract. In the face of escalating global food demand and increasing climate variability, precise and granular crop yield monitoring is indispensable for maintaining regional agricultural stability. However, current deep learning approaches for yield estimation are severely constrained by their heavy reliance on massive in situ labeled data, which limits their application in data-scarce regions. Furthermore, these models often overlook the essential temporal evolution logic of yield formation and lack a systematic discussion regarding the contribution patterns of different feature dimensions, resulting in a black-box nature of the underlying model mechanisms. To address these challenges, this study proposes a field-label-free training framework for maize yield estimation that couples mechanistic model with deep learning. The framework's core strength lies in a physiologically complete simulation database, using the WOFOST model to exhaustively cover 30 years of climate variability and habitat combinations across Northeast China (1.24 × 106 km2). A Gated Recurrent Unit (GRU) network was then introduced for end-to-end modeling, accurately capturing the energy accumulation trajectory from vegetative to reproductive growth. Validation against 458 independent ground points (2022–2024) demonstrated robust generalization with an R2 of 0.69, an RMSE of 1.21 t ha−1, and an RRMSE of 13.73 %, despite using no ground data for training. Our analysis revealed that integrating photosynthetic intensity (LAImean), duration (LAD) and peak features (LAImax) across growth stages is critical for accuracy, while omitting early-stage features significantly impairs the model's ability to capture cumulative growth effects. Furthermore, the model successfully captured the spatiotemporal yield anomalies caused by the 2023 typhoon and flooding events. Ultimately, this study generated a 10 m resolution maize yield dataset (2019–2024) for Northeast China. The dataset exhibits consistent interannual stability, with the RRMSE ranging from 7.98 % to 12.92 % and the R2 remaining above 0.44 at the city level. By deeply coupling mechanistic simulation with data mining, this dataset provides detailed support for optimizing agricultural production and guiding farming practices. The Northeast China Maize Yield 10 m dataset is openly available at https://doi.org/10.5281/zenodo.19547014 (Hu et al., 2026).

Why it matches plant phenotyping methodsトウモロコシ収量という植物・作物群落の形質を推定する計算フレームワークを開発し、独立地点で性能検証したうえで再利用可能な10 m解像度データセットを生成しており、単なる農業実験の routine measurement ではない。

abstractthis study proposes a field-label-free training framework for maize yield estimation that couples mechanistic model with deep learning.
Reproduction assets foundThe paper's core output, the NortheastChinaMaizeYield10m maize yield dataset (2019–2024) with accompanying uncertainty layers, is openly deposited on Zenodo with an explicit availability statement and DOI. No author analysis code or trained model checkpoints are stated as publicly available.
Dataset · publicThe Northeast China Maize Yield 10 m dataset is openly available at https://doi.org/10.5281/zenodo.19547014 (Hu et al., 2026).Open asset ↗Zenodo · 10.5281/zenodo.19547014lines:158-191
Code / dataset availability confirmedEurope PMC · Crossref · checked 15 Sept 2026
Published20 Aug 2026Annals of BotanyCited by 0 · OpenAlex ↗

A modern phytolith reference collection for selected native Australian plants: Implications for vegetation reconstruction

LeafSeed / grainClassification

Background and aims Phytolith analysis is widely applied in palaeoecological and archaeological research, but its interpretive strength depends on the availability of robust modern reference collections. This study expands the modern Australian phytolith reference collection by analysing 42 native plant species representing 24 families and 37 genera with emphasis on silicification patterns across major growth forms, including forbs, shrubs, trees, and C3 grasses. Methods Phytoliths were extracted from available plant parts, including leaves, stems, flowers, seeds, seed pods, cones, and roots, depending on sample availability. Morphotypes were identified following ICPN 2.0, with grass silica short cell phytoliths (GSSCPs) further classified by shape and size to examine subfamily-level patterns. Phytolith morphotype percentage data were analysed using Hellinger transformation, PerMANOVA, PCA, LDA, and hierarchical clustering to assess compositional differences among plant growth forms and grass subfamilies. Key results Phytolith production varied strongly among growth forms and plant parts. Grasses were abundant producers, whereas most forbs, shrubs, and trees were trace producers or non-producers. Leaves were the most consistent source of phytoliths, while seeds and seed pods were predominantly non-producers. Grass silica short-cell phytolith (GSSCP) morphotypes showed clear subfamily-level differentiation. Pooideae produced Rondel morphotypes. Danthonoideae produced Rondel as well as wide Bilobate types. Panicoideae and Oryzoideae exhibited a pronounced Bilobate signature, commonly associated with Polylobate and Cross forms. Non-grass taxa (woody, shrubs, and forbs) were dominated by Spheroids, Tracheary elements, Epidermal, and Polygonal sheets and other non-diagnostic forms. Phytolith assemblages differed significantly among plant families, with Poaceae uniquely producing GSSCPs, while non-grass families showed greater overlap in assemblage composition. Conclusion By expanding taxonomic and anatomical coverage, this study strengthens the capabilities of phytoliths in the reconstruction of grasslands and in general paleo vegetation in Australia, especially where other proxies such as pollen are limited.

Why it matches plant phenotyping methods植物部位の珪酸体を抽出・形態分類し、成長形態やイネ科亜科を識別する現代参照コレクションを構築しており、植物形質の取得・判別手法が研究の中心である。

abstractThis study expands the modern Australian phytolith reference collection by analysing 42 native plant species representing 24 families and 37 genera with emphasis on silicification patterns across major growth forms
Reproduction assets foundThe authors explicitly state that the R scripts used for data analysis and figure generation are publicly available on their GitHub repository, which directly reproduces this paper's phytolith statistical analyses (PCA, LDA, PerMANOVA, clustering, plots). Supplementary data files contain the paper's measurements but no
Code · publicntification of all plant specimens collected for this study. A 14 FUNDING M 15 Funding for this study was provided by ARC Discovery grant DP210100508 and a Ph.D. D 16 fellowship (UQGSS) to MH. TE 17 DATA AVAILABILITY 18 The R scripts used for data analysis and figure generation are publicly available on GitHub EP 19 repository: https://github.com/Manoshi-sporo/Australian-Phytolith-Reference-Collection. 20 CONFLICTS OF INTEREST CC 21 The authors declare no competing financial or commercial interests. A 22 AUTHOR CONTRIBUTIONS 23 MH: writing original draft, conceptualization, software, investigation. AC: Supervision, Writing - 24 Review and editing, FM: Supervision, Writing-Review and editing.Open asset ↗Manoshi-sporo/Australian-Phytolith-Reference-Collectionpdf-layout-page:34 lines:1-87
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 5 Sept 2026
Published12 Aug 2026Plant PhenomicsCited by 0 · OpenAlex ↗

SAM-CLIP-Thermal: Leveraging large multimodal models for reliable and scalable annotation in thermal image segmentation for field plant phenotyping.

Brassica vegetablesField / plotThermalWhole plant / canopy / plot / fieldSegmentationPlant / canopy temperature

Thermal imaging enables non-invasive assessment of canopy temperature, an essential indicator of plant stress, yet the lack of color cues and strong shadow interference make plant segmentation in thermal images difficult. Recent advances in foundation models have demonstrated improved performance and generalizability across applications, showing promise for domain-specific applications with limited annotated datasets such as plant segmentation in thermal images. This study investigates large multimodal models (LMMs) for thermal image segmentation in plant phenotyping. Building upon the SAM-CLIP framework, we design a unified pipeline spanning zero-shot inference, few-shot and low-shot fine-tuning, and active learning to maximize accuracy with minimal supervision. Evaluations on two thermal datasets, LadyBird Brassica and UGA Brassica, demonstrate robust performance after minimal adaptation across both datasets and superior performance compared with baselines, achieving mIoU D values of 97.54% on the LadyBird dataset and 76.94 % on the UGA dataset. We also release the resulting thermal segmentation annotations to support community benchmarking and reproducible research, highlighting the potential of LMMs to enable scalable, high-quality dataset construction for field phenotyping. The released datasets can be found at: https://cornell.box.com/s/dh69xf84464yrc1vlws92l1tflx7qa89

Why it matches plant phenotyping methods熱画像から植物を分割する手法を開発・評価し、植物フェノタイピング用データセットとアノテーションも公開しているため、フェノタイピング手法が中心的である。

abstractThis study investigates large multimodal models (LMMs) for thermal image segmentation in plant phenotyping.
Reproduction assets foundThe authors publicly released the paper-specific thermal segmentation annotations (20,538 LadyBird masks and 37,790 UGA masks) via a Cornell Box link stated in the abstract, results, and data availability statement. No author analysis code or trained model checkpoints are explicitly released; the mmsegmentation GitHub/
Dataset · publicwe generated and publicly released segmentation annotations for the complete LadyBird and UGA thermal image datasets using the best-performing SAM-CLIP model. Specifically, the final model obtained through the multi-round training process was used to generate 20,538 masks for the LadyBird dataset and 37,790 masks for the UGA dataset. Details of the generated annotations are provided in Supplementary Fig. S1 , and both annotated datasets are publicly available at: https://cornell.box.com/s/dh69xf84464yrc1vlws92l1tflx7qa89Open asset ↗lines:220-232
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Aug 2026Scientific dataCited by 1 · OpenAlex ↗

A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science.

ClassificationStress / disease detectionDisease symptoms / severity

Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis. To address this, we present PlantExpertVQA, a large-scale visual question answering (VQA) dataset designed to advance vision-language models for agricultural decision-making. It is compiled from 45 open-source datasets, including the widely used PlantVillage corpus, and comprises 765,186 high-quality question-answer (QA) pairs grounded over 150,841 images spanning 38 crop species and 89 disease conditions. Questions are organized into 3 levels of cognitive complexity and 9 distinct categories. Each was phrased following expert guidance and generated via an automated two-stage pipeline: template-based QA synthesis from image metadata, followed by multi-stage linguistic re-engineering. The dataset was iteratively reviewed by domain experts for scientific accuracy and relevance. We find that current frontier vision-language models, including recent open-source instruction-tuned multimodal LLMs, perform poorly on PlantExpertVQA. However, parameter-efficient fine-tuning of a compact 2B-parameter model on a small fraction of the dataset yields substantial improvements across all question categories, demonstrating its effectiveness for domain adaptation.

Why it matches plant phenotyping methods植物病害画像を対象とする大規模VQAデータセットの構築・ベンチマークが研究の中心であり、植物の病害状態を画像から評価する再利用可能なデータセットです。

abstractwe present PlantExpertVQA, a large-scale visual question answering (VQA) dataset designed to advance vision-language models for agricultural decision-making.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicThe code for the programmatic QA generation pipeline, the data-refinement and template-paraphrasing steps, the automated outlier-detection pipeline, and the parameter-efficient fine-tuning experiments reported in this work is publicly available at https://github.com/syed-nazmus-sakib/PlantExpertVQA.Open asset ↗syed-nazmus-sakib/PlantExpertVQAhtml-lines:578-597
Code / dataset availability confirmedarXiv · OpenAlex · checked 15 Sept 2026
Published3 Aug 2026arXivCited by 0 · OpenAlex ↗

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

MaizeSoybeanWheatAerial / UAVRGB / grayscaleWhole plant / canopy / plot / field2D/3D reconstructionPlant / canopy height

Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. We introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys. It contains 88,830 RGB images at $5280 \times 3956$ pixels, with a ground sampling distance of 3.6-5.8 mm, from 91 scenes spanning corn, soybean, wheat, and oat. Track A evaluates seven scene-optimized methods -- Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) variants -- on held-out views, photogrammetry-referenced depth, and canopy-height recovery. Track B tests four pretrained feed-forward models on zero-shot camera-pose and geometry estimation. The scene-optimized methods rank differently across the three targets: Splatfacto-big leads appearance, whereas Scaffold-GS leads depth and is statistically tied with Splatfacto for canopy height. Among feed-forward models, MapAnything leads on seven of the eight metrics, while the remaining models vary more across crops and fail severely on absolute scale in a way that alignment conceals. Repeated acquisitions reveal further sensitivities that differ by output type and by model, associated with position within the acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/

Why it matches plant phenotyping methods植物キャノピー高さという明示的な形質を対象に、UAV 3D再構成手法をベンチマークし、公開データセットとして提供しているため、フェノタイピング手法が中心である。

abstractWe introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys.
Reproduction assets foundThe paper introduces UAV3DCrop, a public benchmark of repeated multi-angle UAV crop surveys (88,830 RGB images, 91 scenes, four crops) with refined poses, photogrammetric depth references, and linked canopy-height and effective-LAI field measurements. The dataset is explicitly stated to be publicly available under CC B
Dataset · publiche acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/ . Keywords: UAV imagery; agricultural datasets; crop-field reconstruction; neural radiance fields; Gaussian splatting; feed-forward geometry. 1 IntroductionOpen asset ↗UAV3DCroplines:1-90
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published27 Jul 2026Frontiers in Fungal BiologyCited by 0 · OpenAlex ↗

AgriFusionNet: a context-aware multimodal leaf disease diagnosis and classification system for sustainable plant health monitoring

Growth chamberMultimodalLeafClassificationStress / disease detectionDisease symptoms / severity

Early and accurate identification of plant diseases is essential for improving crop productivity and ensuring food security. Many existing deep learning-based plant disease classification methods rely solely on leaf images collected from a controlled environment, which limits their applicability in real-world agricultural conditions where symptoms may be visually unclear and influenced by environmental factors. To address these challenges, this study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification. The proposed architecture employs EfficientNet-B0 for visual feature extraction, BERT for semantic representation of symptom descriptions, and a lightweight multilayer perceptron for modeling environmental factors such as temperature, humidity, rainfall, and soil moisture. Features from all three modalities are fused into a unified representation to train the CNN model. The model is trained and tested upon the Context-Aware Multimodal Augmented PlantVillage dataset covering 38 plant diseases and healthy classes. Experimental results show that AgriFusionNet gives an overall accuracy of 98.94% on the dataset Context-Aware Multimodal Augmented PlantVillage, with competitive precision and recall and F1-score. The multimodal framework facilitates the co-learning of visual, semantic, and contextual environmental representations and the analyses of the confusion matrix and feature interactions give insights into cross-modal relationships. The proposed approach aims to explore context-aware multimodal representation learning for agricultural AI applications, with emphasis on integrating complementary visual, semantic, and contextual information.

Why it matches plant phenotyping methods葉画像を中心に、症状記述と環境情報を統合して植物病害状態を分類する手法を開発・評価しており、植物フェノタイピング手法が中心である。

abstractthis study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification.
Reproduction assets foundThe paper's data availability statement points to the Context-Aware Multimodal Augmented PlantVillage dataset (leaf images, symptom text, environmental data used for the phenotyping/classification analysis) deposited publicly on IEEE Dataport with a DOI matching an allowed URL.
Dataset · publicPublicly available datasets were analyzed in this study. This data can be found here: Dataset. IEEE Dataport. https://dx.doi.org/10.21227/9jat-r836 [Accessed on August 2025].Open asset ↗IEEE Dataport · 10.21227/9jat-r836lines:1029-1047
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published24 Jul 2026Data in briefCited by 0 · OpenAlex ↗

Image dataset of manalagi apple fruits for multi-class disease classification using deep learning.

AppleField / plotRGB / grayscaleFruitClassificationDisease symptoms / severity

This dataset contains images of Manalagi apple diseases from Indonesia. Data collection was conducted from August 2024 to June 2026. Data were collected in apple orchards. All images were captured under natural environmental conditions. A total of 1168 unique Manalagi apple specimens were successfully documented. The specimens consisted of both healthy and diseased fruit. This dataset comprises four classes: Healthy, Anthracnose, Black Pox, and Powdery Mildew. Each specimen was observed and photographed directly. The documentation process yielded approximately 5100 raw images. The images were captured using various smartphone cameras and DSLR cameras. Each device has different camera specifications. The image size depends on the device used. Images that passed quality inspection were selected for the next stage. Each fruit specimen is cropped from the selected raw image. Each image was then labeled according to its disease class. The image size was standardized to 1024 × 1024 pixels. All images were saved in JPEG format. The curation process yielded 482 images. Each image represents a distinct fruit specimen.

Why it matches plant phenotyping methodsリンゴ果実の健全・病害状態を画像で記録し、分類用データセットとして構築・キュレーションした研究であり、植物病害表現型の取得方法と再利用可能なデータセットが中心です。

titleImage dataset of manalagi apple fruits for multi-class disease classification using deep learning.
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository of Manalagi apple fruit disease images (raw, curated, and augmented), directly usable for plant disease phenotyping/classification.
Dataset · publicRepository name: Mendeley Data Data identification number: DOI: 10.17632/9zgkwwv9j8.6 Direct URL to data: https://data.mendeley.com/datasets/9zgkwwv9j8/6Open asset ↗Mendeley Data · 10.17632/9zgkwwv9j8.6html-lines:97-124
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published23 Jul 2026Scientific reportsCited by 0 · OpenAlex ↗

Explainable hybrid multi-branch CNN-ViT-GNN framework for robust hibiscus leaf disease classification.

Field / plotLeafClassificationDisease symptoms / severity

Early and reliable diagnosis of hibiscus leaf diseases is critical to protect horticultural yield. Yet, it remains challenging under real-time field conditions where uncontrolled lighting, clutter, and the non-contiguous nature of pathological symptoms blur diagnostic cues. To address these challenges, we introduce CNN-FusionViT-GNN. This explainable hybrid multi-branch framework synergizes the fine-grained texture extraction of a DenseNet201 backbone, the global contextual modeling of a Vision Transformer (ViT), and the relational reasoning of a Graph Neural Network (GNN). The model is trained and validated on 'Hibiscus,' a curated field dataset of 1165 images from Bangladesh, which is strategically augmented to 8000 samples for robust training following a strict train-validation-test split. The proposed framework achieves a state-of-the-art accuracy of 98.33% with a macro F1-score of 0.98. The framework's generalization is confirmed through high performance on external datasets: 98.78% accuracy on the 52-class Plant City dataset and 83.88% on the 10-class Tomato Leaf Disease dataset, while maintaining a rapid inference time of 10-45 ms. Furthermore, a multi-faceted Explainable AI (XAI) audit using LIME, Grad-CAM++, ViT Attention Maps, and Occlusion Sensitivity validates that the model's decisions are driven by biologically meaningful symptom patterns rather than background artifacts. This study establishes a computationally efficient, transparent, and robust pathway for automated disease diagnosis in precision agriculture.

Why it matches plant phenotyping methodsハイビスカス葉の病徴を画像から分類するCNN-ViT-GNN手法を開発し、複数データセットで性能検証しているため、植物病害表現型の取得・推定が中心である。

abstractwe introduce CNN-FusionViT-GNN. This explainable hybrid multi-branch framework synergizes the fine-grained texture extraction of a DenseNet201 backbone, the global contextual modeling of a Vision Transformer (ViT), and the relational reasoning of a Graph Neural Network (GNN).
Reproduction assets foundThe paper's primary Hibiscus leaf disease image dataset is publicly deposited on Mendeley Data, and the external Tomato Leaf Disease dataset used for validation is also publicly available on Mendeley Data. No author analysis code or trained model checkpoints are reported.
Dataset · publicThe primary dataset generated and analyzed during the current study,“Hibiscus Leaf Diseases Classification Dataset,”is publicly available in Mendeley Data 7 .Open asset ↗Mendeley Datalines:307-347
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published23 Jul 2026Journal of Intelligent Decision Making and Information ScienceCited by 0 · OpenAlex ↗

Multi-Crop Leaf Disease Detection using YOLOv12 with Class-Aware Multi-Scale Fusion and Adaptive Attention Modules

AppleMaizeMangoPotatoSugarcaneTomatoField / plotLeafWhole plant / canopy / plot / fieldObject detection

- Enhancing agricultural productivity and attaining sustainable crop management depend on the early and precise identification of leaf disease. Using state-of-the-art technologies in precision agriculture like machine learning (ML) and image processing greatly increases the effectiveness of disease detection and facilitates well-informed decision-making. But conventional manual inspection techniques are still tedious, unpredictable, and prone to errors. In order to overcome these constraints, this research offers YOLOv12-CropNet, an innovative deep learning-based system for multi-crop leaf disease diagnosis in real time. The proposed YOLOv12-CropNet approach makes use of the Convolutional Block Attention Module (CBAM) for adaptive attention, the Content-Aware Reassembly of Features (CARAFE) up-sampling module to preserve fine-grained disease characteristics, the YOLOv12 architecture improved with Ghost Convolution for effective feature extraction, and Involution layers to capture spatially specific patterns. Inspection techniques are still laborious, arbitrary, and prone to mistakes. A substantial set of data of 38 classes of both healthy and sick leaves from a variety of crops, including tomato, potato, apple, grape, corn, mango and sugarcane, was put together for training and evaluation. Experimental results show that YOLOv12-CropNet finds a suitable balance between computational speed and accurate detection. Accuracy, F1-score, recall, and precision are important performance metrics that verify the model's resilience in challenging environmental and visual circumstances. The suggested technique provides a scalable and field-deployable way to assist effective identification of diseases and precision agricultural decision-making. The proposed YOLOv12-CropNet model exhibits better performance than the other evaluated models, attaining a 98.45% peak accuracy, 98.10% precision ,98.20 % sensitivity and a 98.18% F1 score, thereby highlighting its efficacy in multi-crop leaf disease detection.

Why it matches plant phenotyping methods複数作物の葉の病徴を画像から検出・分類する深層学習手法を開発し、データセットと性能評価を伴うため、植物病害状態のフェノタイピング手法が中心である。

abstractthis research offers YOLOv12-CropNet, an innovative deep learning-based system for multi-crop leaf disease diagnosis in real time.
Reproduction assets foundThe paper uses public Kaggle datasets as its phenotyping image inputs: the PlantVillage dataset (38 crop-disease classes) and the Sugarcane Leaf Disease dataset, both cited with explicit public URLs. No author code, models, or checkpoints are reported as publicly available.
Dataset · public[37] PlantVillage Dataset. Available online: https://www.kaggle.com/datasets/abdallahalidev/plantvillage-datasetOpen asset ↗pdf-page:22 lines:1-61
Dataset · public[39] Sugarcane leaf Disease Dataset available online: https://www.kaggle.com/datasets/nirmalsankalana/sugarcane-Open asset ↗pdf-page:22 lines:1-61
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published20 Jul 2026Scientific dataCited by 0 · OpenAlex ↗

A forty-four-year dataset of rapeseed phenology in the Middle and Lower Yangtze River Plain of China.

Rapeseed / canolaField / plotWhole plant / canopy / plot / fieldAnnotation / quality controlGrowth / time-series analysisGrowth / development / phenology

This study compiles and releases the first standardized rapeseed phenology observation dataset spanning forty-four years (1981-2024) over the core winter rapeseed production region of the Middle and Lower Yangtze River Plain in China. The data originate from systematic observations at 50 national-level agrometeorological stations across six provinces: Jiangsu, Zhejiang, Anhui, Jiangxi, Hubei, and Hunan. The dataset provides complete records of the specific dates for each phenology stage from sowing to maturity, including eight key phenology periods: Sowing (SO), Emergence (EM), Five-leaf (FV), Bud Formation (BF), Stem Elongation (SE), Flowering (FL), Green Ripening (GR), and Maturity (MA), along with the calculated durations of six distinct growth lengths. We implemented a multi-level quality control protocol encompassing internal logical checks, statistical outlier detection, climatological validation, time series homogenization, and expert arbitration. This protocol effectively constrained data uncertainty and corrected non-climatic discontinuities. Univariate linear regression was further employed to quantify the decadal change trends of each phenology period and growth length, supplemented by Kernel Density Estimation (KDE) to characterize their probability distribution features. The final dataset is presented as structured tables (in xlsx format) and high-resolution diagnostic plots (including trend and density plots), with a total volume of approximately 470 MB, systematically organized by province and station. This dataset fills a critical gap in long-term, standardized rapeseed phenology data for the region. The integrated analysis of phenology dates, growth stage durations, and their trends across the entire network provides an indispensable, high-quality empirical foundation. It is designed to support in-depth investigations into the nonlinear response mechanisms of overwintering crops to climate warming, improve crop model parameterization and validation, and inform regional adaptive management strategies.

Why it matches plant phenotyping methods44年間のナタネの生育段階日を標準化・品質管理して公開するデータセット研究であり、植物状態(フェノロジー)の測定データ整備が中心です。

abstractThis study compiles and releases the first standardized rapeseed phenology observation dataset spanning forty-four years (1981-2024)
Reproduction assets foundThe paper's rapeseed phenology dataset (1981–2024, 50 stations) is openly deposited in Science Data Bank under DOI 10.57760/sciencedb.34086, containing structured xlsx tables and diagnostic plots. No custom code was created per the authors.
Dataset · publicThe dataset described in this work has been deposited in the Science Data Bank (ScienceDB) under accession code https://doi.org/10.57760/sciencedb.34086 [27].Open asset ↗Science Data Bank · 10.57760/sciencedb.34086pdf-page:12 lines:1-68
Code / dataset availability confirmedOpenAlex · Europe PMC · Crossref · checked 5 Sept 2026
Published10 Jul 2026Frontiers in Plant ScienceCited by 0 · OpenAlex ↗

SPVD-field: a task-oriented multi-task visual dataset for sweet potato virus disease under real field conditions

PotatoSweet potatoField / plotWhole plant / canopy / plot / fieldAnnotation / quality controlClassificationObject detectionSegmentationStress / disease detectionDisease symptoms / severity

Sweet potato virus disease (SPVD) is one of the most destructive diseases affecting sweet potato production worldwide, causing severe yield losses and posing a significant threat to food security. Vision-based intelligent diagnosis has emerged as a promising solution for large-scale SPVD monitoring due to its low cost and scalability. However, existing publicly available datasets for SPVD are extremely limited and typically focus on a single task, such as disease classification or lesion segmentation, under constrained imaging conditions. This lack of comprehensive, task-oriented datasets significantly restricts the development, evaluation, and fair comparison of advanced computer vision methods for SPVD analysis. In this study, we present SPVD-Field, a task-oriented multi-task visual dataset suite composed of two independently collected sub-datasets optimized for different computer vision tasks. Rather than constructing a single homogeneous dataset, SPVD-Field is deliberately organized into two complementary task-oriented sub-datasets: SPVD-DET, designed for disease detection with bounding-box annotations, and SPVD-SEG, designed for fine-grained lesion segmentation with pixel-level masks. The two sub-datasets were independently collected using different acquisition protocols optimized for their respective tasks, while sharing a unified semantic definition of SPVD symptoms, crop growth stages, and field environments. SPVD-Field captures substantial real-world variability in imaging scale, viewpoint, illumination, background complexity, and symptom manifestation, reflecting the inherent challenges of fieldbased disease diagnosis. We provide detailed documentation of data acquisition, annotation strategies, and quality control procedures, along with baseline benchmark results for both detection and segmentation tasks to demonstrate the usability and difficulty of the dataset. By offering a structured dataset suite rather than a single-task collection, SPVD-Field aims to support diverse research directions, including detection, segmentation, multi-task learning, and disease severity analysis, and to facilitate reproducible and comparable research in SPVD-related plant phenotyping.

Why it matches plant phenotyping methodsサツマイモの病徴を対象とする画像データセットで、検出・病斑セグメンテーション、データ取得・アノテーション・品質管理、ベンチマークを中心的に提供しており、植物病害状態の画像フェノタイピング手法・データ基盤に該当する。

abstractIn this study, we present SPVD-Field, a task-oriented multi-task visual dataset suite composed of two independently collected sub-datasets optimized for different computer vision tasks.
Reproduction assets foundThe paper's core asset is the SPVD-Field dataset (SPVD-DET detection images with bounding-box annotations and SPVD-SEG segmentation images with pixel-level masks), explicitly deposited in a public repository via the data availability statement with a DOI link.
Dataset · publicThe datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: https://dx.doi.org/10.21227/hq1q-jp43 .Open asset ↗10.21227/hq1q-jp43lines:664-703
Code / dataset availability confirmedEurope PMC · Crossref · checked 5 Sept 2026
Published9 Jul 2026Springer Science and Business Media LLCCited by 0 · OpenAlex ↗

A Novel Multi Class Real World Fruit and Leaf Disease Image Dataset for Crop Health Analysis

Pepper / chilliTomatoField / plotRGB / grayscaleFruitLeafWhole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

Abstract Plant diseases affecting leaves and fruits cause substantial yield and economic losses worldwide, particularly in horticultural crops cultivated under diverse agro-climatic conditions. Early and accurate disease diagnosis is essential for effective crop management; however, manual inspection is time-consuming, subjective, and often infeasible at large scale. In this work, we present the Tomato–Chilli–Papaya (TCP) Fruit and Leaf Disease Dataset, a comprehensive multi-crop image dataset designed to support deep learning-based plant disease recognition. The dataset comprises labeled RGB images of healthy and diseased leaves and fruits from three economically important crops—tomato, chilli, and papaya—captured under real-field and semi-controlled environments, reflecting significant variability in illumination, background complexity, and disease severity. To demonstrate the applicability of the dataset, several commonly used convolutional neural network (CNN) architectures, including VGG, ResNet, DenseNet, MobileNet, and EfficientNet models, were trained and evaluated on the TCP dataset using transfer learning. Experimental results show that deep CNN models can effectively learn discriminative visual features corresponding to disease-specific patterns such as leaf spots, lesions, discoloration, curling, and fruit surface abnormalities. Lightweight models such as MobileNet achieve competitive performance with reduced computational cost, while deeper architectures provide improved accuracy at the expense of higher complexity. The results highlight the importance of dataset diversity for robust model generalization across multiple crops and plant organs. The TCP dataset provides a challenging benchmark for single-crop and multi-crop disease classification and supports the development of advanced deep learning, attention-based, and explainable AI models for precision agriculture. By enabling reproducible research and realistic performance evaluation, this dataset contributes toward scalable and practical AI-driven plant disease diagnosis systems aimed at reducing yield losses and supporting sustainable agriculture.

Why it matches plant phenotyping methods植物の葉・果実の病徴を画像から評価する大規模データセットとベンチマークを中心に扱っており、植物病害状態の画像ベース表現型解析に該当する。

abstractwe present the Tomato–Chilli–Papaya (TCP) Fruit and Leaf Disease Dataset, a comprehensive multi-crop image dataset designed to support deep learning-based plant disease recognition.
Reproduction assets foundThe paper introduces the TCP (Tomato-Chilli-Papaya) fruit and leaf disease image dataset and reports CNN experiments on it. The dataset is publicly deposited on Mendeley Data, and the authors state that analysis code is available on GitHub. Both are paper-specific, public, and actionable.
Dataset · publicData is available on Mendeley:1Open asset ↗pdf-page:27 lines:1-51
Code · publicCode availability: Code is available on GitHub 2Open asset ↗GitHubpdf-page:27 lines:1-51
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published3 Jul 2026Plant phenomics (Washington, D.C.)Cited by 0 · OpenAlex ↗

ELMERF: A deep-learning-assisted hydroponic RGB phenotyping framework for rice seedling salt-stress evaluation and genetic mapping.

RiceGrowth chamberRGB / grayscaleRootSegmentationPigment / colour / senescenceStress response / tolerance

Rice seedling salt-tolerance evaluation commonly relies on visual scoring or destructive assays, which are subjective, labor-intensive, and difficult to standardize for population-level analysis. This study developed a new deep-learning-assisted hydroponic RGB phenotyping framework for standardized salt-stress evaluation and genetic mapping in rice seedlings. The framework integrates controlled hydroponic cultivation, RGB imaging, RicePhenoSeg-assisted annotation and trait extraction, ELMERF-based semantic segmentation, and image-derived quantification of salt-induced shoot injury. Using this framework, we constructed the Rice Seedling-Salt RGB Dataset (RSSD), which contains green shoot tissues, yellow shoot tissues, roots, and background from hydroponically grown rice seedlings. Based on RSSD, ELMERF achieved a mean Intersection over Union of 51.4% and a mean Accuracy of 89.5%, outperforming nine representative segmentation models. We further defined shoot yellowing rate (SYR) as an image-derived quantitative trait describing visible salt-induced shoot injury. The framework was applied to 261 re-sequenced rice accessions for population-level phenotyping and genome-wide association analysis. Compared with standard evaluation score and seedling death rate, SYR showed a more continuous phenotypic distribution and detected 36 significant SNPs, including a major signal near the Saltol/OsHKT1; 5 region. Notably, 34 SYR-associated SNPs were not detected by conventional visual scores. Overall, this study provides a targeted hydroponic RGB phenotyping framework for standardized rice seedling salt-stress evaluation and genetic analysis.

Why it matches plant phenotyping methods深層学習によるRGB画像セグメンテーション、形質抽出、データセット構築、性能比較を中核とし、画像由来の塩ストレス傷害形質を定量化する植物フェノタイピング手法である。

abstractThis study developed a new deep-learning-assisted hydroponic RGB phenotyping framework for standardized salt-stress evaluation and genetic mapping in rice seedlings.
Reproduction assets foundThe paper's Data Availability statement explicitly deposits datasets, source code, and supporting data in a public GitHub repository (ELMERF), which covers the RSSD RGB image dataset, segmentation code, and phenotyping/GWAS analysis assets. RiceVarMap is a cited external SNP database, not a paper-specific asset.
Code · publicThe datasets, source code, and other supporting data are openly available on the ELMERF repository (https://github.com/PhenoCodexh/ELMERF).Open asset ↗PhenoCodexh/ELMERFhtml-lines:446-478
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 Jul 2026BMC plant biologyCited by 0 · OpenAlex ↗

PD-ViCo: an explainable AI-based contrastive captioner vision transformer with patch dropout for multi-class brinjal disease classification.

Eggplant / aubergineField / plotFruitClassificationDisease symptoms / severity

Brinjal (eggplant) is a critical crop in South Asia, especially in Bangladesh, but its production is drastically affected by numerous diseases that inhibit yield and quality. Manual diagnosis of disease is time-consuming, subjective, and prone to errors, necessitating automated, scalable technology. To address these issues, this paper proposes PD-ViCo, a lightweight, efficient transformer-based model for brinjal fruit disease classification using Simple Vision Transformer (ViT) with Patch Dropout and Contrastive Captioner (CoCa) methods. One new dataset of 1,823 field-harvested brinjal images encompassing five disease classes including Phomopsis Blight, Fruit and Shoot Borer, Fruit Cracking, Wet Rot, and Healthy samples were prepared through real-world agricultural data collection from Bangladesh. The approach includes extensive preprocessing, class balancing (under-sampling/oversampling), and resilient augmentation methods. The PD-ViCo model significantly improves classification performance under data imbalance with patch dropout regularization and CoCa-style aggregation, resulting in better generalization and robustness. On a range of imbalanced, under-sampled, and oversampled datasets, PD-ViCo achieved a classification accuracy of 99.12% and F1-score of 97.76%, outperforming both ViT and Swin Transformer across all key evaluation metrics. Explainability was also applied using Grad-CAM and Grad-CAM + + , generating visual explanations of model decisions and maintaining conformity to disease-affected regions in the images. These visualizations ensure the credibility of the model and its usability for real agricultural conditions. This study demonstrates that PD-ViCo is a highly accurate, interpretable, and lightweight model for multi-class brinjal disease diagnosis. Not only does it advance state-of-the-art in agricultural AI, but it also provides a valuable dataset and an understandable decision-making protocol that can be applied directly by farmers, agronomists, and agricultural technologists.

Why it matches plant phenotyping methods植物画像から病害状態を分類するモデル、データセット、説明可能性評価を中心に開発・検証しており、植物フェノタイピング手法として適格。

abstractthis paper proposes PD-ViCo, a lightweight, efficient transformer-based model for brinjal fruit disease classification
Reproduction assets foundThe paper's own field-harvested brinjal disease image dataset (1,823 images, five classes) is publicly deposited on Mendeley Data, with an explicit availability statement and URL matching an allowed entry. No code or model checkpoint deposit is stated.
Dataset · publicThe data utilized in this study is publicly accessible on Mendeley Data Repository at the following link: [ https://data.mendeley.com/datasets/ngc58fsxgd/1 ].Open asset ↗Mendeley Data · ngc58fsxgd/1lines:226-251
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published1 Jul 2026Cited by 0 · OpenAlex ↗

TopoLeaf: A Zero-Shot Visual Anomaly Detection Framework via Stable Feature Dimensionality and Local Persistent Homology for Self-Organizing Agricultural Cyber-Physical Systems

AppleMaizeStrawberryStress / disease detectionDisease symptoms / severity

Abstract Zero-shot visual anomaly detection in complex textured domains remains a fundamental challenge for building adaptive, self-organizing cyber-physical systems. Conventional deep learning approaches often rely on closed-set assumptions, require prohibitive pixel-level annotation costs, and suffer severe performance degradation under cross-domain shifts---limiting their deployability in real-world agricultural CPS where novel disease types and unseen crop species continuously emerge. To address these issues, we present TopoLeaf, a training-free and annotation-free anomaly detection framework. By leveraging the robust semantic representations of foundation models (specifically DINOv2), our method introduces two complementary scoring mechanisms: a geometric anomaly score based on local KNN distance in a stability-selected feature subspace, and a topological anomaly score derived from local persistent homology. The topological score effectively captures subtle structural deviations and micro-texture mutations that geometric distances often miss. Extensive experiments on cross-species plant disease benchmarks (3,100+ images across 40 source--target pairs) demonstrate that TopoLeaf achieves highly competitive and structurally robust zero-shot performance, providing a robust perception layer for closed-loop agricultural cyber-physical systems that must maintain diagnostic stability under previously unseen perturbations. Under well-aligned domains, our geometric score achieves near-perfect detection (e.g., 0.994 AUROC on Strawberry). The method exhibits informative failure modes on structurally isolated domains such as Corn (0.169 AUROC), revealing fundamental structural properties of the foundation model's feature manifold. Furthermore, the topological score demonstrates structural complementarity, achieving 0.542 AUROC on the challenging Corn-to-Apple pair where geometric scoring degenerates to 0.221. Module ablation studies confirm that stability-based dimensionality selection consistently improves cross-domain generalization.

Why it matches plant phenotyping methods植物病害の視覚的異常(植物の病徴・状態)を推定する新規画像解析手法を開発し、複数種の病害ベンチマークで検証しているため、植物フェノタイピング手法が中心である。

abstractwe present TopoLeaf, a training-free and annotation-free anomaly detection framework.
Reproduction assets foundThe paper publicly releases its complete TopoLeaf source code (implementation, baselines, evaluation scripts) under the MIT License on GitHub, and all experimental image data derives from the publicly available PlantVillage dataset, which is the leaf-image input used for the paper's anomaly-detection phenotyping and is
Code · public427 7.4 Consent to Publish 428 Not applicable. 429 7.5 Data Availability 430 All experimental data used in this study is derived from the publicly available 431 PlantVillage dataset [17], which can be accessed at https://github.com/spMohanty/ 432 PlantVillage-Dataset. 433 7.6 Code Availability 434 The complete source code, including implementation of TopoLeaf, baseline com- 435 parisons, and evaluation scripts, is publicly available at https://github.com/ 436 Shutong-Hou/TopoLeaf under the MIT License. 437 7.7 Funding 438 This research received no specificOpen asset ↗Shutong-Hou/TopoLeafpdf-layout-page:26 lines:1-44
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published1 Jul 2026Data in briefCited by 0 · OpenAlex ↗

BanglaRiceLeaf: A benchmark dataset for automated rice leaf disease detection and health classification in Bangladesh.

RiceField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Rice leaf diseases pose a major challenge to crop health and agricultural productivity, particularly when timely and accurate diagnosis is required under natural field conditions. The development of automated disease recognition systems depends heavily on the availability of large, well-annotated image datasets. However, many existing rice leaf disease datasets are limited in terms of environmental variability, disease representation, and real-field imaging conditions. To address this gap, this paper presents BanglaRiceLeaf, an original rice leaf image dataset collected and curated by the authors from the experimental fields of the Bangladesh Rice Research Institute (BRRI), Gazipur, Bangladesh, between July 2023 and July 2024. The dataset contains 4152 images belonging to five classes: Bacterial Leaf Blight, Bacterial Leaf Streak, Sheath Blight, Leaf Blast, and Healthy Leaf. The images were acquired from two rice varieties, BR11 and BRRI dhan34, under natural field conditions across varying illumination environments in order to reflect practical disease recognition scenarios. All images were manually annotated by trained annotators under expert supervision. The dataset is systematically organized and publicly released to support reproducible research in rice disease classification. In addition to dataset presentation, benchmark experiments using Xception, NASNetMobile, and InceptionV3 are provided to demonstrate its applicability for deep learning-based disease recognition. BanglaRiceLeaf is expected to serve as a useful resource for plant disease analysis, comparative model evaluation, and future research in precision agriculture and agricultural computer vision.

Why it matches plant phenotyping methodsイネ葉の病徴・健全状態を画像で分類する公開ベンチマークデータセットであり、データ収集・注釈・ベンチマーク評価が中心です。

abstractthis paper presents BanglaRiceLeaf, an original rice leaf image dataset collected and curated by the authors
Reproduction assets foundThe paper's core asset is the BanglaRiceLeaf rice leaf disease image dataset (4152 field images, five classes), publicly released on Harvard Dataverse with DOI 10.7910/DVN/XAOBYW. No author analysis code or trained model checkpoints are stated as publicly available.
Dataset · publicData Identification Number: https://doi.org/10.7910/DVN/XAOBYW Direct URL to Data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/XAOBYW Access Instructions: This dataset is publicly available on the Harvard Dataverse repository and can be accessed for academic, research, and instructional purposes.Open asset ↗Harvard Dataverse · doi:10.7910/DVN/XAOBYWhtml-lines:100-131
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published26 Jun 2026Earth System Science DataCited by 0 · OpenAlex ↗

CropPlantHarvest: a 500 m annual dataset of crop planting and harvesting dates (2001–2024) of the U.S. Midwest

MaizeSoybeanField / plotGreenhouseMultispectral / hyperspectralWhole plant / canopy / plot / fieldGrowth / time-series analysisTrackingGrowth / development / phenologyYield / yield components

Abstract. As key components of agricultural management, planting and harvesting schedules have strongly influenced crop production by defining the length of the crop growing season and shaping the environmental conditions crops experience. Accurate knowledge of these management data is crucial for enhancing crop yield estimates by capturing the timing of crop development relative to weather and soil conditions, assessing climate adaptation by tracking shifts in farming practices over time, and supporting agricultural carbon accounting. Yet, existing planting and harvesting date datasets are largely based on state-level statistics or rule-based calendars that overlook intra-regional variability and the influence of human decision-making. The absence of long-term, high-resolution planting and harvesting date information hinders our ability to reconstruct historical agricultural practices and assess their agronomic and environmental consequences. In this study, we introduce CropPlantHarvest, the first dataset of annual corn and soybean planting and harvesting dates across the U.S. Midwest at 500 m resolution from 2001 to 2024. Planting dates are estimated using CropSow, an integrative remotely sensed crop modeling system that aligns simulated crop growth trajectories with satellite observations to retrieve field-level planting dates. Harvesting dates are retrieved using the Normalized Harvest Phenology Index (NHPI), a novel index that integrates Normalized Difference Vegetation Index (NDVI) and near-infrared (NIR) reflectance to detect harvesting events by capturing the distinct spectral transition from senescent crops to exposed crop residues. Validation against USDA crop progress reports and field-level dataset demonstrates high accuracy of CropPlantHarvest, with a mean absolute error of approximately 5 d for both crop species. This large spatial and temporal dataset captures management-driven variability in crop season timing and duration, supporting improved modeling of crop yields, greenhouse gas emissions, and resource use. It could also serve as a benchmark for refining remote-sensing phenology products and evaluating the agro-environmental impacts of evolving crop management decisions. CropPlantHarvest is available at https://doi.org/10.5281/zenodo.16967482 (Liu and Diao, 2025).

Why it matches plant phenotyping methods衛星観測と作物モデルによる圃場レベルの作付・収穫時期推定手法を開発し、NHPIを提案して独立データで検証した大規模データセット研究であり、植物の生育・収穫状態の取得が中心的です。

abstractPlanting dates are estimated using CropSow, an integrative remotely sensed crop modeling system that aligns simulated crop growth trajectories with satellite observations to retrieve field-level planting dates.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicOur CropPlantHarvest dataset, which provides planting and harvesting dates for corn and soybean fields at 500 m spatial resolution across the U.S. Midwest from 2001 to 2024, can be accessed via Zenodo: https://doi.org/10.5281/zenodo.16967482 (Liu and Diao, 2025).Open asset ↗Zenodo · 10.5281/zenodo.16967482lines:322-333
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published26 Jun 2026Scientific reportsCited by 0 · OpenAlex ↗

Multi-modal deep learning for paddy health assessment: fusing leaf imagery with tabular metadata using a factorized bilinear pooling approach.

RiceLeafClassificationDisease symptoms / severity

Global food security is largely based on the accurate and timely diagnosis of crop diseases, where paddy rice is an extremely essential staple of more than half of the world population. The conventional disease identification techniques tend to be laborious, time consuming and demand a great deal of domain knowledge, which becomes a bottleneck in the efficient management of the farms. Although deep learning [and especially Convolutional Neural Networks (CNNs)] have demonstrated a spectacular performance in automated classification of diseases based on leaf images, they tend to overlook important contextual features that are implicitly processed by agronomic experts. The visual defects of a disease might be unclear and this can greatly differ depending on factors like the genetic variety of the plant and the stage of development. We overcome this shortcoming by proposing a new multi-modal deep learning framework, Multi-Modal Factorized Bilinear Pooling (MFBP) model which is capable of a more holistic and precise paddy health measurement. The proposed method is the only one that combines high-level visual information obtained using leaf images and related tabular information, namely the paddy type and number of days. The MFBP model uses Factorized Bilinear Pooling (FBP) rather than the simple feature concatenation which commonly loses the complex relationship between different data types. This systematic method efficiently encodes all the complex interactions between all components of the visual and tabular features vectors in such a way that helps the model to pick up subtle, context-specific patterns. As an example, it will only be possible to educate the model that a specific visual blemish is predictive of a given disease through a specific species at a specific age. We test our model on the Paddy Doctor: Paddy Disease Classification dataset, which is a detailed public dataset comprising of more than 10,000 labeled images and containing relevant metadata, and thus it forms a perfect testing bed to conduct multi-modal research. Through our detailed experiments, we have shown that the proposed MFBP model is much better than a baseline model based on concatenation fusion, which proves that deep, multiplicative interactions can be best modeled in this task. The findings highlight the massive possibilities of multi-modes AI in the development of more robust, more accurate, and more context-aware diagnostic instruments and precision agriculture to enable more sustainable and productive agricultural activities.

Why it matches plant phenotyping methods葉画像とメタデータを統合してイネの健康状態・病害を推定する新規深層学習手法を提案し、ベースライン比較で検証しているため、植物フェノタイピング手法が中心である。

abstractWe overcome this shortcoming by proposing a new multi-modal deep learning framework, Multi-Modal Factorized Bilinear Pooling (MFBP) model which is capable of a more holistic and precise paddy health measurement.
Reproduction assets foundThe paper's phenotyping inputs are the public Kaggle 'Paddy Doctor: Paddy Disease Classification' dataset (10,407 leaf images with tabular metadata for variety and age), explicitly named in the Data Availability statement with a persistent public URL. No author analysis code, trained models, or checkpoints are reported
Dataset · publicThe datasets used and/or analysed during the current study are publicly available in the “Paddy-doctor: paddy disease classification” repository at the following persistent URL: https://www.kaggle.com/datasets/vbookshelf/paddy-disease-classification.Open asset ↗Kaggle · vbookshelf/paddy-disease-classificationpdf-page:20 lines:1-74
Code / dataset availability confirmedEurope PMC · Crossref · checked 5 Sept 2026
Published23 Jun 2026Plant phenomics (Washington, D.C.)Cited by 0 · OpenAlex ↗

A-Occ-Plant: Plant occluded point cloud completion via amodal segmentation

SoybeanField / plotNeRF / 3D Gaussian SplattingLiDAR / point cloudLeafWhole plant / canopy / plot / fieldMorphology / geometry measurementPose / keypoint estimation2D/3D reconstructionSegmentation

Plants are geometrically and topologically complex objects, and methods and devices that produce plant point clouds often miss parts due to self occlusions, making further analysis, such as phenotypic trait extraction or 3D reconstruction, difficult. We introduce A-Occ-Plant , a novel method for point cloud completion. The first novelty of our algorithm is converting point clouds into a set of images, which are then completed using 2D amodal segmentation. The images are then converted into a complete point cloud by using view-consistent Gaussian splats. The second novelty is the use of a coarse-to-fine hierarchical Transformer with cross-scale attention. The completed soft masks are fused into a continuous 3D density field using Gaussian splatting, removing the need for external pose estimation or fixed-size inputs. We introduce a synthetic dataset using a procedural model and a real-world plant reconstruction benchmark with artificially generated occlusions. We further benchmark A-Occ-Plant against representative 3D point-cloud completion methods, demonstrate that it recovers downstream phenotypic traits (leaf count, leaf angle, plant height), and show that it generalizes to another crops (soybean). A-Occ-Plant achieves a 264.8% improvement in LPIPS and an 8.3% gain in SSIM compared to the current state of the art, while using only 2.3% of the parameters and running 39.4× faster. We release our code at https://github.com/JaeLee18/PlantPhenomics_Occlusion.

Why it matches plant phenotyping methods植物の遮蔽点群を補完し、葉数・葉角度・草丈という表現型形質を復元する手法を開発しており、データセット作成とベンチマーク検証も中心的に行っている。

abstractWe introduce A-Occ-Plant , a novel method for point cloud completion.
Reproduction assets foundThe paper explicitly releases authors' code and sample data (inference code, sample data for reproducing results) via a Google Drive project download and a GitHub repository, both with explicit availability statements and public URLs.
Code · publicThe full code and data at https://github.com/JaeLee18/PlantPhenomics_Occlusion .Open asset ↗JaeLee18/PlantPhenomics_Occlusionlines:386-410
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published23 Jun 2026Plants (Basel, Switzerland)Cited by 0 · OpenAlex ↗

Rapid Classification and Deep Learning-Based Development Estimation of the Seeds of Helianthus annuus .

SunflowerLaboratory / benchtopSeed / grainClassificationCountingObject detectionFruit / seed / panicle traits

Manually counting sunflower seeds on capitula is labor-intensive, requiring approximately one person-hour per head, and can be inconsistent for densely packed heads. Existing phenotyping approaches often depend on laboratory-based equipment, limiting their accessibility. In this study, we developed a benchtop image-based pipeline for rapid, non-destructive estimation of developed and aborted seeds on intact dried sunflower heads. A dataset of 1093 sunflower capitula was imaged under fixed indoor lighting, and individual seeds were annotated as developed or aborted. A YOLOv8m one-stage object detector was trained and evaluated using a counting-focused protocol, in which a single confidence threshold was selected on the validation set and then applied unchanged to an independent test set of 109 images. The baseline model was compared with recent YOLO variants and different augmentation strategies. On the test set, the model achieved a mean absolute count error of 61.3 seeds per image, a mean relative error of 12.0%, and an mAP50 of 0.18 at the locked confidence threshold of 0.15. Only 13.8% of test images had relative errors below 2%. Larger YOLO models and augmentation variants did not improve performance. These findings show that the proposed system provides approximate, non-destructive seed-count estimation under controlled imaging conditions, while highlighting the need for improved localization in dense regions and domain adaptation for fresh heads or field conditions. The annotated dataset and trained model weights are made available to support reproducible research.

Why it matches plant phenotyping methodsヒマワリ頭花の発達・不稔種子数という植物形質を、画像取得とYOLOによる推定パイプラインで定量化する手法を開発・評価しており、方法が研究の中心である。

abstractwe developed a benchtop image-based pipeline for rapid, non-destructive estimation of developed and aborted seeds on intact dried sunflower heads.
Reproduction assets foundThe authors state the source code is available on GitHub and the CVAT-annotated dataset is available via a public share link; the GitHub repository URL is explicitly provided and matches an allowed URL. The dataset link itself is not given, so only the code/checkpoint repository qualifies as an actionable public asset.
Code · publicThe developed system is available as a Telegram bot [ 19 ] and the source code is available on GitHub [ 20 ]. The CVAT annotated dataset is available via a public share link.Open asset ↗lines:84-103
Code / dataset availability confirmedOpenAlex · Crossref · checked 15 Sept 2026
Published23 Jun 2026DataCited by 0 · OpenAlex ↗

LeafScans-Orchard: A Multi-Year Open RGB Scan Dataset of Orchard Plant Leaves for Species and Cultivar Classification

AppleCherryPeachPearPlumLaboratory / benchtopRGB / grayscaleLeafClassificationMorphology / geometry measurement

LeafScans-Orchard is a curated, multi-year RGB image dataset of orchard plant leaves designed to support research in computer vision, machine learning, and plant phenotyping. The dataset comprises 9708 high-quality leaf scans acquired during collection campaigns conducted between 2015 and 2025, covering seven orchard crop species: apple, pear, sweet cherry, sour cherry, plum, peach, and apricot. In total, the dataset includes 67 cultivar labels. All samples were acquired using flatbed scanning under controlled conditions on a uniform background, ensuring high visual consistency and minimal background variability. The original scans were captured at 1200 dpi and subsequently converted into a public release format at 300 dpi, stored as lossless TIFF images to preserve morphological and textural details. Each image corresponds to a single leaf and is organized in a hierarchical directory structure by species, cultivar, and acquisition year, accompanied by image-level metadata and aggregated species–cultivar–year counts. LeafScans-Orchard is suitable for plant species classification, cultivar recognition, leaf morphology analysis, texture analysis, and general visual feature extraction. In addition to the main release, a representative subset of 300 original 1200 dpi scans is provided to support high-resolution analyses. The dataset is particularly suited for fine-grained classification, morphology-driven analysis, and methodological studies under controlled imaging conditions.

Why it matches plant phenotyping methods果樹葉のRGBスキャン画像を収録した公開データセットで、植物フェノタイピングおよび葉形態解析を目的とする。標準化された画像取得と再利用可能なデータ構成が中心であり、フェノタイピング用データセットとして適格。

abstractLeafScans-Orchard is a curated, multi-year RGB image dataset of orchard plant leaves designed to support research in computer vision, machine learning, and plant phenotyping.
Reproduction assets foundThe paper's core asset is the LeafScans-Orchard dataset itself (9708 RGB leaf scans, 300 dpi TIFF release plus 1200 dpi subset, image-level metadata and summary counts), openly deposited on Zenodo with an explicit DOI and CC BY 4.0 license. This is a paper-specific, public, actionable phenotyping image dataset. No code
Dataset · publicthe published version of the manuscript. Funding: This research received no external funding. Institutional Review Board Statement: Not applicable. Informed Consent Statement: Not applicable. Data Availability Statement: The dataset described in this article is openly available in Zenodo as LeafScans-Orchard Dataset (v1.0.0) at https://doi.org/10.5281/zenodo.20187966 (accessed on 10 May 2026). The repository includes the 300 dpi image release, the 1200 dpi high-resolution subset, image-level metadata, aggregated species–cultivar–year counts, and supporting documentation. The complete archive of original 1200 dpi scans is retained locally by the authors but is not included in the current pubOpen asset ↗Zenodo · 10.5281/zenodo.20187966pdf-raw-page:12 lines:1-46
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published23 Jun 2026Scientific dataCited by 0 · OpenAlex ↗

RoseVisuals: A Multi-Class Dutch Rose Petal Images Dataset for Automated Health and Pigmentation Classification via Deep Learning.

FlowerClassificationDisease symptoms / severityPigment / colour / senescence

The robust Dutch rose, also known as the Rosa hybrida is distinguished by its vibrant colors, superior product quality, and extended vase life. These rose varieties, originating from Netherlands, have proven highly successful in Indian agricultural conditions and the international export industry. The dataset consists of a total of 1,995 high resolution petal image collected during this research, encompassing petal color categories, such as red, yellow, white, pink, purple, orange, bi-color, and multi-color, as well as health statuses including fresh, dry, and diseased petals. The primary purpose of this dataset is to support machine learning activities in agriculture and specifically for tasks such as automatic petal health evaluation and rose variety categorization. Although the rose flower is scientifically rich and has a wide range of industrial uses, it has not been given much attention in machine learning, especially when compared to other plant-based datasets. This study adds to the accuracy of quality assessment through the use of modern computer vision and machine learning methods, thus helping the agriculture sector, rose-based edible product making, and flavor development industries.

Why it matches plant phenotyping methodsバラ花弁画像データセットの構築と、花弁の健康状態・色分類による植物状態評価が研究の中心であり、画像ベースの表現型計測データセットに該当する。

abstractThe dataset consists of a total of 1,995 high resolution petal image collected during this research, encompassing petal color categories, such as red, yellow, white, pink, purple, orange, bi-color, and multi-color, as well as health statuses including fresh, dry, and diseased petals.
Reproduction assets foundThe paper's own rose petal image dataset is publicly deposited on Mendeley Data, and the authors' validation/metadata scripts are publicly available on GitHub. Both are paper-specific, public, and actionable.
Dataset · publicThe RoseVisuals dataset is publicly available on Mendeley Data at Direct URL to data: https://data.mendeley.com/datasets/f44jwtbfjg/5. Data Identification Number: 10.17632/f44jwtbfjg.5. Repository Name: RoseVisuals.Open asset ↗Mendeley Data · 10.17632/f44jwtbfjg.5html-lines:246-284
Code · publicThe RoseVisuals codebase, comprising all validation scripts, is publicly available on GitHub Repository at https://github.com/Arya-S14/RoseVisuals-Validation-Doc.Open asset ↗GitHubhtml-lines:246-284
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published21 Jun 2026Data in briefCited by 0 · OpenAlex ↗

A curated image dataset for sapodilla fruit (Manilkara zapota) disease and fruit quality analysis.

Field / plotRGB / grayscaleFruitClassificationStress / disease detectionDisease symptoms / severity

Sapodilla (Manilkara zapota), or chickoo, is a key tropical fruit, very popular in India, Mexico, and Thailand, as it is nutritionally and economically valuable. Nonetheless, the production of sapodillas is often affected by several diseases, which reduce fruit quality and quantity. The dataset used in this paper is a sapodilla fruit image dataset, comprising 1,518 images, gathered in the field under the practicing conditions on 18 February 2025, 22 February 2025, in Rahu village, Pune district, Maharashtra, India, with the use of smartphone cameras. The data is sorted into four categories, namely: Anthracnose, Bacterial rot, Healthy, and Sap bleeding. The photographs were taken in different backgrounds and in different lighting conditions to represent real-life cultivation conditions. The data is expected to be useful in machine learning-based plant disease detection, classification, and analysis, and spur the creation of intelligent and sustainable agricultural systems.

Why it matches plant phenotyping methodsサポディラ果実の病害・健全状態を画像で記録したデータセット自体が中心で、植物病害の画像ベース表現型解析に利用できる。

titleA curated image dataset for sapodilla fruit (Manilkara zapota) disease and fruit quality analysis.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicData identification number: Version: V1, Doi: 10.17632/xbzd2fjd3p.1 Direct URL to data: https://data.mendeley.com/datasets/xbzd2fjd3p/1Open asset ↗10.17632/xbzd2fjd3p.1html-lines:1-114
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published20 Jun 2026Plant methodsCited by 0 · OpenAlex ↗

Automated stomatal traits measurement in melon (Cucumis melo L.) based on vision transformers with dynamically composable multi-head attention.

MelonStomata / guard-cell complexMorphology / geometry measurementSegmentationStomatal traits

Stomatal trait analysis is essential for optimizing crop photosynthesis and transpiration, yet deep learning studies have focused mainly on monocotyledons, leaving dicotyledonous crops such as melon (Cucumis melo L.) understudied. To bridge this gap, we established a dedicated melon stomatal dataset comprising 5,708 training images, 1,631 validation images, and 815 test images. On this basis, we developed an improved Mask R-CNN framework using Vision Transformer (ViT) as the backbone. Specifically, standard Multi-Head Attention (MHA) was replaced with Dynamically Composable Multi-Head Attention (DCMHA), which enhances information exchange across attention heads and alleviates the low-rank limitation of conventional attention. In addition, a modified effective Squeeze-and-Excitation (eSE) module was incorporated into the Feature Pyramid Network (FPN) to strengthen channel dependency modeling and multi-scale feature representation. On the melon dataset, the proposed model achieved a mean average precision (mAP) of 72.40 ± 0.09%, with AP50 and AP75 of 91.93 ± 0.14% and 84.59 ± 0.19%, respectively. Repeated-run statistical analyses showed that eSE significantly and consistently improved the main detection metrics across backbones, whereas DCMHA provided a more moderate gain within the ViT-based setting, with clearer support for AP50 than for mAP or AP75 under the baseline FPN setting. Overall, the combined configuration remained among the top-performing models for stomatal instance segmentation. Ellipse fitting further enabled automated quantification of stomatal length, width, count, area, and circumference, showing strong agreement with manual measurements (Pearson r = 0.978). The model also showed preliminary transferability to cucumber, watermelon, pumpkin, and loofah, with an average species-specific R² of 0.86, although each species was evaluated on a limited sample set.

Why it matches plant phenotyping methodsメロンの気孔形質を画像から自動抽出するデータセット、改良Mask R-CNN、インスタンスセグメンテーション、楕円フィッティングを開発・検証しており、植物フェノタイピング手法が研究の中心である。

abstractwe established a dedicated melon stomatal dataset comprising 5,708 training images, 1,631 validation images, and 815 test images.
Reproduction assets foundThe paper's authors explicitly state that the source code for model training and inference (including the key modules: DCMHA, eSE-enhanced FPN, Mask R-CNN/ViT pipeline) is publicly available at a GitHub repository, which matches an allowed URL. The melon stomatal image dataset (8,154 images) is described in detail but,
Code · publicCode Availability The source code for model training and inference, including the implementation of the key modules, is publicly available at: https://github.com/huangyao110/qk_maskrcnn_trsv2.gitOpen asset ↗huangyao110/qk_maskrcnn_trsv2pdf-page:22 lines:1-322
Code / dataset availability confirmedOpenAlex · arXiv · checked 5 Sept 2026
Published16 Jun 2026arXiv (Cornell University)Cited by 0 · OpenAlex ↗

Vines-DB: An RGB image dataset for multi-species ornamental vine segmentation

Field / plotRGB / grayscaleWhole plant / canopy / plot / fieldSegmentation

The Vines-DB dataset contains 1,218 original high-resolution RGB images of seven ornamental vine species collected under field conditions at the Utah Agricultural Experiment Station's Greenville Research Farm in Logan, Utah, USA. The dataset was generated from 168 individual vine plants that were transplanted in 2022 and photographed repeatedly across multiple months during the 2023 and 2024 growing seasons (July-October). Images were captured with an iPhone 16 Pro equipped with a 48 MP camera between 10:00 AM and 12:00 PM under daylight. Vines were grown on 1.2m x 2.4m trellises and photographed from a distance of 1m against black or white Styrofoam backdrops to improve contrast and reduce background noise. The dataset includes Akebia quinata, Campsis radicans, Hydrangea anomala petiolaris, Lonicera x heckrottii, Campsis x tagliabuana 'Madame Galen', Parthenocissus quinquefolia, and Wisteria floribunda. All original images were manually annotated in Roboflow by trained annotators to produce polygon-based instance segmentation masks for eight classes, including seven species and background. After preprocessing and data augmentation, the working dataset was expanded to 2,307 images for model development and evaluation. The augmented dataset was divided into 2,019 training images, 192 validation images, and 96 test images using stratified sampling to maintain balanced representation. Vines-DB supports the development and evaluation of deep learning models for multi-class instance segmentation in precision horticulture and urban ecology. The dataset enables applications such as automated canopy cover estimation, species identification, and scalable field phenotyping. In addition, repeated monthly imaging of the plants captures temporal variation in canopy development and plant appearance, increasing the dataset's utility for segmentation benchmarking under realistic field conditions.

Why it matches plant phenotyping methods植物のRGB画像とポリゴン注釈から成るデータセットを構築し、セグメンテーション評価およびキャノピー被覆推定などの植物フェノタイピングを支援することが中心であるため。

abstractVines-DB supports the development and evaluation of deep learning models for multi-class instance segmentation in precision horticulture and urban ecology.
Reproduction assets foundThe paper's core asset is the Vines-DB RGB image dataset with instance segmentation annotations, publicly deposited on OSF with an explicit DOI and URL matching an allowed URL.
Dataset · publicData accessibility Repository name: Vines-DB Data identification number: 10.17605/OSF.IO/YJHCK Direct URL to data: https://osf.io/yjhck/overviewOpen asset ↗OSF · 10.17605/OSF.IO/YJHCKpdf-page:2 lines:1-49
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published15 Jun 2026Scientific reportsCited by 0 · OpenAlex ↗

Wheat spike and spikelet detection and counting from high-resolution digital imagery using YOLO with Oriented Bounding Boxes.

WheatRGB / grayscalePanicle / ear / spikeCountingObject detectionFruit / seed / panicle traits

In-season estimation of wheat grain yield potential is critical for crop management and advancing breeding efforts. Spike and spikelet counts serve as key indicators directly linked to yield potential, yet their assessment still relies on manual counting which is both labor-intensive and error-prone. High-resolution digital (RGB) imagery combined with deep learning-based object detection methods has substantially advanced automatic wheat spike detection and counting. However, precise spikelet-level phenotyping remains largely underexplored. This study evaluates two recent YOLO variants, YOLOv11 and YOLOv12, for wheat spike and spikelet detection and counting using oriented bounding boxes (OBB), and introduces a new large-scale benchmark dataset comprising 48,521 spike and 60,404 spikelet instances with OBB annotations. For spike detection, the pre-trained YOLOv11 achieved superior accuracy (mAP@0.5 = 95.8%, Pearson r = 0.993) with shorter training and inference times compared to YOLOv12. For spikelet detection, the non-pretrained YOLOv11 demonstrated higher accuracy (mAP@0.5 = 99.0%), while counting performance was comparable across models. These results establish OBB-based YOLO detection as a robust and scalable approach for AI-driven wheat phenotyping.

Why it matches plant phenotyping methods小麦の穂・小穂という収量関連形質の画像ベース検出・計数手法を比較評価し、大規模ベンチマークデータセットも構築しているため、フェノタイピング手法が中心である。

abstractThis study evaluates two recent YOLO variants, YOLOv11 and YOLOv12, for wheat spike and spikelet detection and counting using oriented bounding boxes (OBB), and introduces a new large-scale benchmark dataset comprising 48,521 spike and 60,404 spikelet instances with OBB annotations.
Reproduction assets foundThe paper openly states its supporting data (spike/spikelet imagery with OBB annotations) is available on Zenodo, and the underlying models are deployed on the authors' public WheatAI cloud platform.
Dataset · publicData availability The data supporting the findings of this study are openly available at: https://doi.org/10.5281/zenodo.20215489 .Open asset ↗zenodo · 10.5281/zenodo.20215489lines:219-266
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 15 Sept 2026
Published12 Jun 2026Plant PhenomicsCited by 0 · OpenAlex ↗

Deep learning-driven automatic counting of petal number in cut chrysanthemum inflorescence.

FlowerPanicle / ear / spikeCountingFruit / seed / panicle traits

The number of petals in an inflorescence is an important phenotypic indicator for quality evaluation and cultivar identification of cut chrysanthemums ( Chrysanthemum morifolium Ramat.). Current manual measurement methods are time-consuming, error-prone, and poorly suited to the complex geometry of chrysanthemum flowers, which limits their utility for large-scale phenotyping and breeding programs. Although image-based phenotyping has advanced rapidly, automated and reliable methods for petal counting in densely packed or partially obscured inflorescences remain underdeveloped. Here, we developed a deep learning-based framework for automatic extraction of petal number in cut chrysanthemums. Images from multiple varieties were collected to construct a representative dataset, and petal density maps were generated through manual annotation with Gaussian kernel function. We employed a Congested Scene Recognition Network (CSRNet) enhanced with a Squeeze-and-Excitation (SE) channel attention mechanism (SE-CSRNet) for petal density estimation. Spearman correlation analysis revealed strong agreement between visible and actual petal counts (Spearman’s r=0.953, p<0.0001). Compared with the original CSRNet, SE-CSRNet reduced mean absolute error (MAE) and root mean squared error (RMSE) by 5.2% and 7.4%, respectively. Further optimization using regression fitting revealed that random forest achieved the best performance (MAE = 4.24, RMSE = 5.06, R 2 = 0.967), indicating reliable stability and satisfactory generalization under the conditions evaluated in this work. Application of the optimized model to two cut chrysanthemum varieties confirmed its practicality by successfully detecting reductions in petal number under high-temperature stress. Our results demonstrate that integrating dataset construction, deep learning–based density estimation, and machine learning optimization enables efficient and accurate prediction of petal number in cut chrysanthemums.

Why it matches plant phenotyping methods花弁数という植物形質を画像から自動抽出する深層学習手法を開発し、データセット構築、性能比較、検証、実用適用まで行っており、表現型取得手法が研究の中心である。

abstractHere, we developed a deep learning-based framework for automatic extraction of petal number in cut chrysanthemums.
Reproduction assets foundThe article states that some data (the chrysanthemum petal-counting dataset and related materials) will be available at the authors' public GitHub repository (qwsdfgz/petalscount), with other data available from the corresponding author upon reasonable request. The repository URL is explicitly provided by the authors,但
Dataset · publicnctional components of bud-leaves and flowers in edible chrysanthemum (Chrysanthemum morifolium Ramat) Horticulturae 11 5 2025 448 10.3390/horticulturae11050448 Appendix A Supplementary data The following is the Supplementary data to this article. Multimedia component 1 Data availability Some data will be available at this URL: https://github.com/qwsdfgz/petalscount . Other data are openly available from the corresponding author upon reasonable request. Appendix A Supplementary data to this article can be found online at https://doi.org/10.1016/j.plaphe.2026.100238 .Open asset ↗qwsdfgz/petalscountlines:602-636
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published10 Jun 2026Discover foodCited by 0 · OpenAlex ↗

Toward accurate prediction of apple firmness and brix across countries, seasons and cultivars with hyperspectral imaging.

AppleMultispectral / hyperspectralFruitPhysiological trait estimationFruit / seed / panicle traits

Traditional apple maturity assessment methods are destructive and time- and labour-intensive, yielding only population-level approximations. Hyperspectral imaging provides a non-destructive alternative to assess individual fruit, but progress has been constrained by the lack of large, diverse datasets that support robust model generalisation. This study presents a multi-cultivar, multi-season, multi-country hyperspectral apple dataset to enable generalisable prediction of soluble solids content (Brix) and firmness. Using this dataset, we adopt an iterative modelling framework to evaluate deep learning architectures, image resolutions, cultivar encoding, seasonal effects, and feature-specific models. Wavelength and spatial region importance were also analysed. The best predictive performance was achieved using Vision Transformer (ViT) models trained on edge-cropped 40 × 40 pixel images with explicit cultivar encoding, with Brix and firmness modelled independently. Although seasonal specificity was observed, models trained across all three seasons achieved the strongest overall performance. A 50% reduction in spectral wavebands did not compromise prediction accuracy. Key wavelength ranges contributing to Brix and firmness prediction were identified across the visible-near-infrared spectrum. Spatial regions were unimportant for Brix prediction but showed relevance for firmness. The optimised ViT model achieved firmness prediction performance comparable to previous studies (RMSE = 0.76 kgf, R[Formula: see text] = 0.63), while Brix prediction accuracy was lower (RMSE = 0.91 [Formula: see text]Brix, R[Formula: see text] = 0.75), likely reflecting increased biological and environmental variability captured in the dataset. Overall, this work demonstrates that hyperspectral imaging combined with deep learning and large, diverse datasets enables robust, non-destructive prediction of apple quality attributes across production conditions.

Why it matches plant phenotyping methodsリンゴ果実の硬度とBrixという植物器官形質を、ハイパースペクトル画像と深層学習で非破壊推定するデータセット・モデル・汎化性能評価が研究の中心である。

abstractThis study presents a multi-cultivar, multi-season, multi-country hyperspectral apple dataset to enable generalisable prediction of soluble solids content (Brix) and firmness.
Reproduction assets foundThe paper explicitly states that the hyperspectral apple dataset (5756 apples, firmness/Brix/starch measurements) is deposited in the University of Essex research data repository and that the data cleaning, model training, and analysis code is on GitHub, both with public URLs.
Dataset · publicThe datasets generated during and analysed during the current study are available in the University of Essex repository ( https://researchdata.essex.ac.uk/228/ )Open asset ↗researchdata.essex.ac.uk · 228lines:192-220
Code · publicthe code used for data cleaning, model training and analysis are available on GitHub: ( https://github.com/EIS-Ressearch-Lab/Apple_maturity_hyperspectral_imaging.git )Open asset ↗github.com/EIS-Ressearch-Lab/Apple_maturity_hyperspectral_imaginglines:192-220
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 15 Sept 2026
Published6 Jun 2026Plant PhenomicsCited by 0 · OpenAlex ↗

A knowledge-driven unified framework for plant disease classification and severity grading via domain adaptation.

LeafClassificationStress / disease detectionDisease symptoms / severity

Plant leaf disease classification and severity grading are essential for precision agriculture, enabling timely intervention and optimized management. Existing models often fail to recognize previously unseen disease categories due to rigid label spaces and limited representation of plant phenotypes. To address these challenges, we propose a knowledge-driven unified framework for plant disease classification and severity grading. A Common Knowledge Learner consolidates fundamental features of plant species, disease categories, and severity levels from labeled data, forming a transferable representation space. It employs a multi-level contrastive learning strategy to capture both global semantic representations and fine-grained lesion patterns. Building on these representations, a Cross-Domain Adaptation module leverages a teacher-student framework with Low-Rank Adaptation (LoRA) bridges in-domain and out-of-domain feature spaces using large-scale unlabeled data. Meanwhile, a contrastive feature library enables similarity-based reasoning and supports flexible label space expansion during inference without retraining. We evaluate our approach on Leaf-CG, a large-scale dataset comprising 441,448 images from 59 plant species, 373 disease categories, and four severity levels. Experiments demonstrate that our framework outperforms existing baselines, achieving 94.9% disease classification accuracy and 90.6% severity grading accuracy in-domain. Under out-of-domain conditions, the method achieves 82.1% true positive rate (TPR) in open-set settings, highlighting its strong generalization ability and potential for practical plant disease management. Code and dataset are available at https://www.uniplantcg.samlab.cn.

Why it matches plant phenotyping methods植物画像から病害分類と病害重症度を推定する計算・画像ベース手法を開発し、大規模データセットで評価しており、フェノタイピング手法が中心である。

abstractwe propose a knowledge-driven unified framework for plant disease classification and severity grading.
Reproduction assets foundThe paper's Leaf-CG dataset (test subset publicly available), analysis code, and trained model weights (plant.pth, disease.pth, severity.pth) are explicitly released at the authors' site https://www.uniplantcg.samlab.cn. Cited datasets (AI Challenger 2018, PlantVillage, etc.) are prior work, not paper-specific assets.
Code · publicving 94.9% disease classification accuracy and 90.6% severity grading accuracy in-domain. Under out-of-domain conditions, the method achieves 82.1% true positive rate (TPR) in open-set settings, highlighting its strong generalization ability and potential for practical plant disease management. Code and dataset are available at https://www.uniplantcg.samlab.cn . Keywords Plant disease diagnosis Knowledge-driven learning Domain adaptation pmc-status-qastatus 0 pmc-status-live yes pmc-status-embargo no pmc-status-released yes pmc-prop-open-access yes pmc-prop-olf no pmc-prop-manuscript no pmc-prop-legally-suppressed no pmc-prop-has-pdf yes pmc-prop-has-supplement yes pmc-prop-pdf-onlyOpen asset ↗uniplantcg.samlab.cnlines:1-29
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published4 Jun 2026Data in briefCited by 0 · OpenAlex ↗

Longitudinal multispectral image dataset for ToBRFV disease detection in tomato and pepper plants.

Pepper / chilliTomatoGreenhouseRGB / grayscaleMultispectral / hyperspectralWhole plant / canopy / plot / fieldStress / disease detectionDisease symptoms / severity

ToBRFV is a major threat to tomato and pepper crops because it spreads quickly and survives for a long time in the environment. Since there are few ways to control it after infection, early detection before symptoms are visible is crucial. Yet, only limited public datasets are available for this research. We present one of the first openly accessible, longitudinal multispectral image dataset dedicated to ToBRFV detection. In this study, two tomato cultivars and two pepper cultivars, all of which are commercially important and widely cultivated in greenhouses, were selected. Using these plants ensures that the dataset reflects real-world agricultural practices and captures variability across commercially grown types. Both healthy and ToBRFV-inoculated plants from each cultivar were included in the imaging process. All plants were cultivated under fully controlled greenhouse conditions in Adana Province, Türkiye. Healthy and infected tomato plants were grown in two separate greenhouses to prevent cross-contamination. Imaging was conducted over a 29-day period using Red-Green-Blue (RGB) and Visible Near Infrared (VNIR) cameras, including narrowband captures at 800 nm and 1000 nm, from multiple viewing angles. Infection status was confirmed via Reverse Transcription quantitative Polymerase Chain Reaction (RT-qPCR) analysis at multiple time points. The dataset is organized into four clean, labelled subsets and released under a CC BY 4.0 license. This resource provides unique opportunities for developing and benchmarking computer vision and machine learning approaches for pre-symptomatic plant disease detection, spectral feature analysis, and integration into precision agriculture systems. By combining controlled experimental design, spectral diversity, and open access, it establishes a robust foundation for cross-disciplinary research in plant pathology, agricultural engineering, and artificial intelligence.

Why it matches plant phenotyping methods植物病害状態を対象にした縦断マルチスペクトル画像データセットであり、公開データセットとして開発・ベンチマーク利用を目的とするため、表現型取得が中心です。

abstractWe present one of the first openly accessible, longitudinal multispectral image dataset dedicated to ToBRFV detection.
Reproduction assets foundThe article is a Data in Brief describing the authors' own openly released longitudinal multispectral plant image dataset (ToBRFV-LMID) for tomato and pepper disease detection, deposited on Zenodo under CC BY 4.0 with a direct DOI URL. This is a paper-specific, public, directly actionable phenotype/image asset. No code
Dataset · publicData accessibility Repository name: ZENODO Data identification number: 10.5281/zenodo.17244968 Direct URL to data: https://doi.org/10.5281/zenodo.17244968Open asset ↗ZENODO · 10.5281/zenodo.17244968html-lines:98-126
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published3 Jun 2026Scientific dataCited by 1 · OpenAlex ↗

A curated dataset of 3,477 high-resolution Grapevine (Vitis vinifera) leaf images for automated detection of Black Rot, Esca, and Leaf Blight diseases.

GrapevineField / plotLeafStress / disease detectionDisease symptoms / severity

We introduce the Grapevine (Vitis vinifera) Leaf Image Dataset (GVLiD), a carefully curated set of 3,477 annotated images of Grapevine (Vitis vinifera) leaves to catalyze research in computer vision and plant pathology. Whereas the PlantVillage and Hermos datasets, for instance, contain mainly scanned or laboratory-acquired leaves, GVLiD features vineyard in situ images along with detailed metadata (GPS, lighting, weather, and device model) and expert-verified annotations. To measure the reliability of the annotation, label consistency was very high (κ = 0.86-0.92; 95% CI) as assessed by inter- and intra-rater agreement. Besides Indian viticulture, the dataset also aims to support the field of foliar disease detection in precision agriculture and ML benchmarking, which face significant challenges due to variable illumination and natural leaf backgrounds under field conditions. GVLiD is intended to enable worldwide, reproducible, real-world testing of AI systems for crop disease monitoring.

Why it matches plant phenotyping methodsブドウ葉の病徴を画像で注釈化したデータセットであり、植物病害状態の画像ベース表現型測定と再現可能なベンチマークが中心です。

abstractWe introduce the Grapevine (Vitis vinifera) Leaf Image Dataset (GVLiD), a carefully curated set of 3,477 annotated images of Grapevine (Vitis vinifera) leaves to catalyze research in computer vision and plant pathology.
Reproduction assets foundThe paper's grapevine leaf image dataset (GVLiD) is deposited on Mendeley Data, but that URL is not among the allowed URLs. The authors' validation/analysis code (image-quality metrics, metadata-completeness checks, annotation-reliability calculations) is publicly available on GitHub at the allowed URL, with explicit '
Code · publicAll validation scripts (image-quality metrics, metadata-completeness checks, and annotation-reliability calculations) are publicly available in the GVLiD GitHub repository. This ensures full reproducibility of all validation results reported here. git clone https://github.com/MilindGayakwad/DNNOpen asset ↗https://github.com/MilindGayakwad/DNNhtml-lines:255-349
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published3 Jun 2026Cited by 0 · OpenAlex ↗

An Attention-Enhanced MobileNetV2 with Squeeze-and-Excitation Architecture for Efficient Potato Leaf Disease Detection and Classification

PotatoLeafClassificationStress / disease detectionDisease symptoms / severity

Abstract Potatoes are one of the major crops eaten in developing countries; however, their production is falling due to various diseases. Early identification and detection of potato leaf diseases play a vital role in improving potato quality and quantity. Existing methods are either computationally resource intensive or lack trust in their decision-making process, which makes them difficult to deploy for real-time potato disease classification and limits its accessibility. To mitigate these limitations, this study proposed an attention-enhanced MobileNetV2 with a squeeze-and-excitation architecture, which balances high accuracy with low computational resources. This method incorporates the strength of MobileNetV2 and Squeeze-and-excitation networks. A total of 2152 images of early blight, late blight, and healthy leafs were obtained from the Kaggle public repository, which are partitioned into 70% training, 20% validation, and 10% testing and were utilized to train, validate, and test the proposed model. The MobileNetV2 backbone is utilized for feature extraction, and then a squeeze-and-attention block is used to recalibrate the feature maps by focusing on important features and suppressing irrelevant ones. Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented to visualize the most relevant region of the leaf for decision-making, which increases model interpretability and user trust. The proposed model achieves a remarkable performance of 99% testing accuracy with 9.41 MB total parameters. The proposed model is suitable for real-time potato leaf disease detection and classification, which can be easily accessible to agricultural stakeholders, including farmers, and contributes to food security.

Why it matches plant phenotyping methodsジャガイモ葉の病害状態を画像から分類する深層学習手法を提案・評価しており、植物表現型(病害状態)の取得・推定が中心である。

abstractthis study proposed an attention-enhanced MobileNetV2 with a squeeze-and-excitation architecture, which balances high accuracy with low computational resources.
Reproduction assets foundThe paper's phenotyping inputs are 2152 potato leaf images (early blight, late blight, healthy) obtained from a public Kaggle repository, explicitly stated as publicly available in the Declarations. No author analysis code is shared (Code Availability: Not applicable), and no trained model checkpoints are released.
Dataset · publicAvailability of Data: The datasets generated during and/or analyzed during the current study are publicly available at https://www.kaggle.com/datasets/faysalmiah1721758/potato-dataset.Open asset ↗Kaggle · faysalmiah1721758/potato-datasetpdf-page:23 lines:1-35
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published1 Jun 2026Data in BriefCited by 1 · OpenAlex ↗

TomatoPGT: A 3D point cloud dataset of tomato plants for segmentation and plant-trait extraction.

TomatoGreenhousePhotogrammetry / SfM / MVSLiDAR / point cloudRGB / grayscaleWhole plant / canopy / plot / fieldMorphology / geometry measurementSegmentationArchitecture / morphology / geometry

Three-dimensional (3D) point-cloud phenotyping enables non-destructive and repeatable characterization of plant architecture, supporting the measurement of traits such as internode length, branching topology, and organ orientation. This article presents TomatoPGT (Tomato Plant Graph Twin) , a 3D tomato dataset designed for research on semantic/instance segmentation, graph-based structural representation, and graph-derived phenotypic trait extraction. The dataset contains 42 scans from three greenhouse-grown tomato plants acquired across early to mid-vegetative development using a rotational multi-view imaging system. Each scan consists of 60-70 overlapping RGB images captured under uniform illumination and reconstructed into a metrically scaled dense colored point cloud using Structure-from-Motion and multi-view stereo. TomatoPGT provides: (i) multi-view RGB images, (ii) dense colored point clouds, (iii) manually curated semantic and instance annotations at organ level, (iv) graph representations encoding plant topology and geometry, and (v) tabulated phenotypic traits computed deterministically from the graphs (internode length, insertion angles, and phyllotactic angles). TomatoPGT supports reproducible development and evaluation of 3D phenotyping pipelines, including learning-based segmentation and graph-based modeling of plant architecture.

Why it matches plant phenotyping methods植物の3D形態表現型抽出を目的としたデータセットで、画像・点群・器官アノテーション・グラフ・形質値を提供し、再現可能なフェノタイピング手法の開発と評価を直接支援している。

abstractThis article presents TomatoPGT (Tomato Plant Graph Twin) , a 3D tomato dataset designed for research on semantic/instance segmentation, graph-based structural representation, and graph-derived phenotypic trait extraction.
Reproduction assets foundThe paper's own TomatoPGT dataset (multi-view RGB images, dense point clouds, semantic/instance annotations, graph representations, and CSV phenotypic traits) is publicly deposited on Mendeley Data, and the authors' Cloud-Seg/Cloud-Graph software tools plus supplementary materials (camera calibrations, example datasets
Dataset · publicRepository name 1: Mendeley[2]. Data identification number: DOI: 10.17632/72md54c7n7.1 Direct URL to data: https://data.mendeley.com/datasets/72md54c7n7/1Open asset ↗Mendeley · 10.17632/72md54c7n7.1html-lines:105-178
Code · public6. Code and documentation: CloudSeg and CloudGraph software tools, environment specifications, and example usage instructions are hosted on Zenodo[3].Open asset ↗Zenodohtml-lines:264-308
Code / dataset availability confirmedCrossref · OpenAlex · checked 14 Sept 2026
Published1 Jun 2026Environmental Research: EcologyCited by 1 · OpenAlex ↗

Ecological insights from transferable plant biomass mapping across the arctic using high-resolution structure-from-motion and LiDAR data

Aerial / UAVField / plotPhotogrammetry / SfM / MVSLiDAR / point cloudRootWhole plant / canopy / plot / fieldObject detectionYield / biomass estimationBiomass / plant weightStress response / tolerance

Abstract Warmer temperatures, permafrost thaw, and increased wildfire activity are driving rapid ecological change across the Arctic, significantly altering plant productivity and aboveground biomass (AGB). These rapid changes highlight the urgent need to improve monitoring of vegetation dynamics in the Earth’s northern ecosystems, where high spatiotemporal heterogeneity occurs at scales finer than those captured by traditional satellite observations. The growing use of unoccupied aerial systems (UASs) presents an opportunity to overcome this limitation. Yet, the diversity of UAS platforms, sensors, and data collection and processing workflows presents challenges for developing standardized, generalizable approaches. To address this challenge, we compiled 672 AGB plots co-located with 183 UAS-based structure-from-motion (SfM) or light detection and ranging (LiDAR) surveys collected across the Arctic. Here, we: (1) evaluated the generalizability of UAS-derived canopy structure derived from high-resolution SfM and LiDAR for estimating AGB, (2) assessed scaling errors and their sources in two recent satellite-based AGB products derived from Landsat and moderate resolution imaging spectroradiometer, and (3) demonstrated the use of high-resolution AGB maps to quantify biomass variation across tundra plant functional types (PFTs) and to monitor post-fire recovery. Our results show that both SfM and LiDAR accurately captured AGB and its variability across tundra PFTs using a random forest model (overall root mean squared error: 0.332 kg m –2 ), with mapping performance varying slightly by region and data source. Using UAS-derived AGB maps as a benchmark, we identified systematic biases in satellite-derived AGB products, largely attributable to the magnitude of AGB and structural heterogeneity within coarse-resolution pixels. Applying our model to repeat UAS surveys following a tundra fire on Seward Peninsula, we observed rapid AGB recovery in non-shrub patches, with biomass recovering to pre-fire levels within two years. In contrast, shrub patches recovered more slowly, with AGB gains continuing over 2–4 years through both in-patch growth and lateral expansion (via dispersal) into remaining burned areas. Overall, these findings support the generalizability of UAS-based SfM and LiDAR data for estimating tundra AGB and highlight the need for broader collection and synthesis of such data to improve ecological monitoring and model benchmarking in the Arctic.

Why it matches plant phenotyping methodsUASのSfMおよびLiDARから植物群落の地上部 biomass (AGB) を推定する手法の一般化性能を評価し、衛星推定値のベンチマークにも用いており、植物形質取得が研究の中心である。

abstractevaluated the generalizability of UAS-derived canopy structure derived from high-resolution SfM and LiDAR for estimating AGB
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicThe codes and training data is available on GitHub: https://github.com/Daryl-Open asset ↗pdf-page:20 lines:1-30
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published30 May 2026Data in briefCited by 0 · OpenAlex ↗

Field-based and close-range multispectral imaging dataset for Huanglongbing (HLB) detection in orange trees: A resource for machine learning and digital agriculture.

CitrusField / plotMultispectral / hyperspectralLeafClassificationStress / disease detectionDisease symptoms / severity

This article presents a multispectral imaging dataset dedicated to training a machine learning algorithm for the in situ detection of Huanglongbing (HLB). HLB, also known as citrus greening disease, is a major pathology caused by the bacterial pathogen Candidatus Liberibacter asiaticus , particularly in species of the citrus genus. The dataset is constituted of terrestrial images acquired in a commercial sweet orange orchard of the variety Pera Rio ( Citrus sinensis (L.) Osbeck). The images describe large portions of canopy, with healthy leaves and sections infected by HLB as well as some confounding factors naturally present in orchards. Multispectral images were acquired with a multi-lens camera within the visible-near-infrared domain, resulting in 14 narrow spectral bands. The image acquisition was conducted during two field campaigns in 2023 and 2024. In total, the dataset contains 2,978 images divided into two classes HLB (1,681) and non-HLB (1,297). Originally, data are stored in TIFF format as 14 monochromatic images, organised by spectra band. Additionally, an HDF5-format version is provided, where images are stored as 3D arrays with spectral bands in ascending order. This format is compatible with various programming languages, enables efficient data handling, and is optimised for machine learning and image processing applications, supporting reproducible and portable analysis. This dataset is a valuable resource for the development and benchmarking of classification models, including deep learning approaches, aimed at the detection of HLB. Phytopathology imaging datasets are scarce yet essential for advancing digital agriculture and the development of robust tools for crop disease detection worldwide.

Why it matches plant phenotyping methods柑橘葉・樹冠のマルチスペクトル画像からHLB感染状態を推定するデータセットであり、植物病害状態の表現型取得と機械学習ベンチマークを中心とする。

abstractThis article presents a multispectral imaging dataset dedicated to training a machine learning algorithm for the in situ detection of Huanglongbing (HLB).
Reproduction assets foundThe article is a Data in Brief describing a public multispectral HLB citrus image dataset deposited on Data INRAE (Recherche Data Gouv, DOI 10.57745/054NAB), plus an authors' GitHub repository with preprocessing, registration, and model training scripts. Both are paper-specific, public, and directly actionable.
Dataset · publicData accessibility Repository name: Data INRAE Data access link: https://doi.org/10.57745/054NABOpen asset ↗Data INRAE · 10.57745/054NABhtml-lines:89-123
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published27 May 2026Scientific reportsCited by 0 · OpenAlex ↗

An uncertainty-aware evaluation framework based on hierarchical vision transformers for robust cross-domain plant leaf disease classification.

Field / plotLaboratory / benchtopLeafClassificationStress / disease detectionDisease symptoms / severity

Plant leaf disease detection is a critical task in precision agriculture, where reliable diagnosis under real-world conditions is essential for reducing crop losses and supporting timely intervention. Although deep learning models have achieved high classification accuracy, their performance often degrades under domain shift between controlled laboratory datasets and real-field environments, while predictive uncertainty and confidence calibration remain largely unaddressed.This study presents an uncertainty-aware cross-domain evaluation framework based on a Hierarchical Vision Transformer (HViT) for plant leaf disease classification. The framework integrates multi-scale feature learning with Monte Carlo Dropout-based predictive uncertainty estimation and temperature-based calibration to systematically analyze model behavior in terms of accuracy, reliability, and robustness. Experiments were conducted on two complementary datasets: the New Plant Diseases Dataset (controlled conditions) and the PlantDoc dataset (field conditions), enabling bidirectional cross-domain evaluation. Results demonstrate that the proposed framework achieves superior performance, attaining 97.8% accuracy on controlled data and 93.6% on field data, while significantly improving calibration with lower Expected Calibration Error (ECE = 0.032 / 0.041), reduced Negative Log-Likelihood, and lower Brier score compared to baseline CNN and transformer models. Furthermore, the framework exhibits improved robustness under domain shift, with reduced performance degradation and stable uncertainty behavior. Overall, this study highlights the importance of integrating uncertainty estimation and calibration within a hierarchical transformer-based framework, providing a more reliable and deployment-ready solution for real-world agricultural disease diagnosis.

Why it matches plant phenotyping methods植物葉の病害状態を直接推定する不確実性-aware分類フレームワークの開発・評価が中心であり、異なる条件のデータセット間で精度、校正、頑健性を検証している。

abstractThis study presents an uncertainty-aware cross-domain evaluation framework based on a Hierarchical Vision Transformer (HViT) for plant leaf disease classification.
Reproduction assets foundThe paper's Data availability statement explicitly links the two public image datasets used for its cross-domain plant leaf disease classification experiments: the New Plant Diseases Dataset on Kaggle and the PlantDoc dataset on Dataset Ninja. No author analysis code, models, or checkpoints are reported as available.
Dataset · publicThe New Plant Diseases Dataset can be obtained from Kaggle at [https://www.kaggle.com/datasets/vipoooool/new-plant-diseases-dataset]Open asset ↗Kaggle · vipoooool/new-plant-diseases-datasetlines:360-398
Dataset · publicThe PlantDoc dataset is available for download at [https://datasetninja.com/plantdoc#download]Open asset ↗plantdoclines:360-398
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published25 May 2026Scientific dataCited by 0 · OpenAlex ↗

A High-Resolution Multifocal RGB Pollen Grain Image Dataset for Deep Learning Computer Vision Tasks from Biobío Region, Chile.

Laboratory / benchtopMicroscopyRGB / grayscaleCell / cellular structureClassificationSegmentation

PollenBB16 is an RGB pollen image dataset of Chilean flora with pixel-accurate instance segmentation masks, whose annotation was fully verified by an expert palynologist to guarantee the taxonomic reliability of every published instance. The dataset is designed to close a concrete gap in existing palynological datasets, which typically combine low taxonomic diversity, few samples per class, and low-resolution crops restricted to bounding boxes. PollenBB16 contains 16,198 brightfield optical microscopy images at the native resolution of 3088 × 2064 pixels and 36,383 pixel-accurate polygons across 16 species from the Biobío Region, spanning endemic, native and exotic species of high ecological and melliferous value such as Eucryphia glutinosa and Quillaja saponaria (endemic), Gevuina avellana and Aristotelia chilensis (native), and Medicago sativa and Brassica rapa (introduced). Each spatial position is recorded at three focal planes. The displacement along the z axis reveals features of the exine together with information on the internal structure of the grain that remain inaccessible on a single plane. From this multifocal information, more robust convolutional networks can be trained with more accurate classification. The operational quality of the dataset is backed by a leakage-safe partition that keeps the three focal planes of the same position in the same subset to avoid metric inflation, complemented by a YOLO11n-seg baseline trained for 50 epochs that reaches 0.985 mask mAP@50 on the validation set, establishing a reproducible reference point. Beyond deep learning, PollenBB16 enables interdisciplinary applications in aerobiology, biodiversity monitoring under climate change, ecological restoration of the South American temperate forest, and botanical-origin authentication of Chilean monofloral honeys.

Why it matches plant phenotyping methods植物由来の花粉粒を対象とした高解像度画像データセットとセグメンテーション基準を構築し、深層学習による画像解析を再現可能な形で検証しているため、植物フェノタイピング手法・データセットとして中心的です。

abstractPollenBB16 is an RGB pollen image dataset of Chilean flora with pixel-accurate instance segmentation masks
Reproduction assets foundThe paper's PollenBB16 pollen image dataset (16,198 multifocal RGB microscopy images with pixel-accurate instance segmentation masks) and the accompanying authors' script polygons_to_bboxes.py are publicly deposited on Zenodo (10.5281/zenodo.19830051), per the Data Availability Statement.
Code · publicThe only custom code distributed with this Data Descriptor is the Python script polygons_to_bboxes.py, which regenerates the YOLO bounding-box labels in labels_bb/ from the polygon labels in labels/. The script is packaged inside the scripts/ folder of the same Zenodo repository that hosts the dataset (10.5281/zenodo.19830051) and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, with no restrictions on access.Open asset ↗Zenodo · 10.5281/zenodo.19830051html-lines:731-797
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published22 May 2026Frontiers in Plant ScienceCited by 1 · OpenAlex ↗

Quantifying the reliability gap in cross-domain plant disease classification: benchmarking the limited efficacy of standard mitigation techniques under controlled-to-field shift

Field / plotLaboratory / benchtopLeafWhole plant / canopy / plot / fieldClassificationObject detectionCalibration / preprocessingStress / disease detectionVisualization / data managementDisease symptoms / severity

Introduction Confidence calibration, selective prediction, out-of-distribution scoring, and deep ensembles are mature techniques in machine learning, yet their efficacy under the severe domain shift encountered when plant disease classifiers move from controlled laboratory imagery to heterogeneous field photographs has not been systematically benchmarked. Methods Models trained on PlantVillage were evaluated on PlantDoc leaf-level crop images under a parent-image-aware split protocol, and a suite of standard mitigation techniques was applied to characterize the reliability gap. Analyses included temperature scaling and selective prediction for a fine-tuned ResNet-50, quantitative image-level shift analysis, Grad-CAM visualization, simple target-aware adaptation baselines, frozen-feature backbone comparisons, and ensemble baselines. Results In the primary case study, a fine-tuned ResNet-50 suffered a 67.7-percentage-point accuracy collapse upon cross-domain transfer, while mean predicted confidence remained at 79.76%. Post-hoc temperature scaling reduced calibrated ECE to 0.3645 but left selective risk at 80% coverage at 64.30%. Quantitative image-level shift analysis confirmed large-effect-size differences in saturation ( d = 3.90), border edge density ( d = 3.33), and foreground-occupancy proxy ( d = 2.48) between the two domains, while Grad-CAM visualizations showed that the model shifts attention from lesion-centered regions in PlantVillage to background-dominated areas in PlantDoc. Simple target-aware mitigations, including adaptive batch normalization and feature moment matching, improved accuracy from 0.321 to 0.343 and 0.366, respectively, whereas DANN-style adversarial adaptation degraded performance to 0.252. A frozen-feature backbone comparison across five backbones showed that, within the energy-scoring frozen-backbone comparison, DINOv2-S/14 achieved the highest unknown-detection AUROC (0.764) and the lowest selective risk at 80% coverage (0.520), with paired Wilcoxon tests confirming statistically significant accuracy and macro-F1 differences across backbones. Two ensemble baselines were evaluated: a warm-start end-to-end ResNet-50 ensemble reduced calibrated ECE to 0.063 but achieved only 0.666 AUROC, while a lightweight DINOv2 linear-probe ensemble achieved 0.779 AUROC after calibration but under limited epistemic diversity. Discussion Neither ensemble established deployment-grade reliability: the best selective risk at 80% coverage across all configurations remained above 0.51. The principal contribution is a reproducible, deployment-oriented reliability characterization showing that standard post-hoc and lightweight adaptation techniques reduce but do not eliminate the severe reliability gap under controlled-to-field transfer in agricultural computer vision.

Why it matches plant phenotyping methods植物病害画像分類の信頼性・ドメインシフト・校正・選択的予測を体系的にベンチマークしており、病害状態を画像から推定する方法の技術評価が中心である。

abstracttheir efficacy under the severe domain shift encountered when plant disease classifiers move from controlled laboratory imagery to heterogeneous field photographs has not been systematically benchmarked.
Reproduction assets found本文中に内容が明示された植物フェノタイピング関連の補足表と、その公開リンクを確認しました。
Supplement · publicSupplementary Table 1 ) was therefore constructed by normalizing all labels to a canonical Crop_Disease format and retaining only those categories for which an unambiguous semantic match existed in both datasets.Open asset ↗lines:335-337
Code / dataset availability confirmedEurope PMC · checked 8 Sept 2026
Published21 May 2026Scientific dataCited by 0 · OpenAlex ↗

A Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis.

Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationCountingGrowth / development / phenology

Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.

Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画データセットであり、自動フェノロジー検出と空間検証のためのベンチマーク基盤が中心である。

abstractThe Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions.
Reproduction assets foundThe paper describes a public multi-modal Actinidia chinensis phenology dataset (annotated images, georeferenced videos, ground-truth counts) deposited on Zenodo, plus authors' MIT-licensed preprocessing scripts on GitHub. CVAT and FiftyOne are generic third-party tools and excluded.
Dataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available at Zenodo: https://doi.org/10.5281/zenodo.17371025.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:12 lines:1-92
Code · publicCustom scripts for dataset preparation are publicly available under the MIT License at https://github.com/Open asset ↗GitHubpdf-page:12 lines:1-92
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published20 May 2026Scientific reportsCited by 1 · OpenAlex ↗

A hybrid approach for citrus disease detection using convolutional neural networks and fuzzy inference systems for enhanced accuracy and interpretability.

CitrusLeafClassificationDisease symptoms / severity

The citrus diseases are affecting the fruit production worldwide thereby posing an economical burden. Major research is moving towards finding solutions using Artificial Intelligence (AI) and Image processing methods. Due to factors like illumination variations, leaf form, and disease symptoms, image data has intrinsic uncertainties that are typically difficult for traditional machine learning techniques to handle. In this paper, the interpretability of fuzzy logic is combined with the resilience of deep learning to propose a novel Fuzzy Convolutional Neural Network (Fuzzy-CNN) architecture for the automated diagnosis of citrus leaf diseases. The hybrid method uses a Convolutional Neural Network (CNN) to obtain complex features of citrus images, and a Fuzzy Inference System (FIS) to improve the classification results. The proposed approach encodes accurate data into fuzzy sets and applies linguistic concepts to determine the severity of a disease, which will contribute to the further development of the decision. In order to test and verify the proposed approach, several experiments were carried out, which proved that Fuzzy-CNN is more effective than regular CNN models with the approximate accuracy difference approximately 1.8, and especially in cases when the symptoms of disease are not clear. To strengthen experimental validation, the proposed method is evaluated on two independent datasets, including an external benchmark dataset, imbalance-aware evaluation metrics are employed to ensure robustness and generalizability. Experimental results demonstrate consistent and statistically significant improvements over existing neuro-fuzzy and machine learning approaches. This research contributes to early detection by collaborating the potential of fuzzy neural networks and offering a flexible solution for real-time disease detection in citrus crops.

Why it matches plant phenotyping methods柑橘葉画像から病害および重症度を推定するFuzzy-CNN手法を開発し、独立データセットとベンチマークで検証しており、植物フェノタイピング手法が中心である。

abstractpropose a novel Fuzzy Convolutional Neural Network (Fuzzy-CNN) architecture for the automated diagnosis of citrus leaf diseases.
Reproduction assets foundThe paper's Data Availability statement links two public image datasets used for the citrus disease phenotyping/classification experiments (a Mendeley citrus leaves dataset and a Kaggle orange fruit dataset), and a third public Kaggle citrus disease dataset is cited as the external benchmark dataset used for validation
Dataset · publicThe data used in the current study is publicly available from the following links. [https://data.mendeley.com/datasets/3f83gxmv57/2]Open asset ↗data.mendeley.com · 3f83gxmv57/2html-lines:525-539
Dataset · publicThe data used in the current study is publicly available from the following links. [https://data.mendeley.com/datasets/3f83gxmv57/2] [https://www.kaggle.com/datasets/sgandhi2003/orange-fruit-dataset]Open asset ↗www.kaggle.com · sgandhi2003/orange-fruit-datasethtml-lines:525-539
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 May 2026Data in briefCited by 0 · OpenAlex ↗

A multi-stage, pixel-level annotated apple dataset for precision agriculture research.

AppleField / plotRGB / grayscaleFruitClassificationObject detectionSegmentationGrowth / development / phenology

This article presents a comprehensive dataset of 1406 RGB images of apples ( Malus domestica ), covering three key growth stages-immature (green), semi-mature (color transition), and mature (red). The dataset serves as a resource for detecting and segmenting apples across different developmental phases. Each image includes pixel-level instance segmentation masks annotated in JSON format using the VGG Image Annotator (VIA), ensuring compatibility with deep learning frameworks. The dataset's real-world variability-spanning lighting conditions, occlusions, and clustered fruit arrangements-enhances its utility for training generalizable computer vision models in precision agriculture. It supports tasks such as fruit detection, segmentation and growth-stage classification, addressing the scarcity of annotated data for transitional maturity phases. With 2574 annotated apple instances, this dataset facilitates research on maturity grading and transfer learning for agricultural robotics. By standardizing annotations and incorporating diverse field conditions, this dataset reduces preprocessing overhead and accelerates the development of deployable AI solutions for orchard management. It is particularly valuable for improving model robustness in heterogeneous environments, thereby advancing data-driven horticultural practices.

Why it matches plant phenotyping methodsリンゴ果実の発育段階・成熟度という植物器官の状態を対象に、画素単位アノテーション付き画像データセットを構築しており、観測・抽出手法の再利用可能な基盤が中心である。

abstractThis article presents a comprehensive dataset of 1406 RGB images of apples ( Malus domestica ), covering three key growth stages-immature (green), semi-mature (color transition), and mature (red).
Reproduction assets foundThe paper is a data descriptor for a public apple image dataset (1406 RGB images, pixel-level instance segmentation masks in JSON) deposited on Mendeley Data with a direct URL and DOI, matching the allowed URL exactly.
Dataset · publicRepository name: Wang, Dandan; Wang, Bo (2026), “A Multi-Stage, Pixel-Level Annotated Apple Dataset for Precision Agriculture Research”, Mendeley Data, V4 Data identification number: 10.17632/gfcmdbvw65.4 Direct URL to data:https://data.mendeley.com/datasets/gfcmdbvw65/4Open asset ↗Mendeley Data · 10.17632/gfcmdbvw65.4html-lines:1-97
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published16 May 2026Data in briefCited by 0 · OpenAlex ↗

Handheld hyperspectral imaging dataset of annual sowthistle and little mallow under abiotic stress for machine learning.

GreenhouseMultispectral / hyperspectralWhole plant / canopy / plot / fieldClassificationCalibration / preprocessingStress response / tolerance

Machine learning has become an increasingly important tool for overcoming agricultural challenges by enabling efficient and consistent classification of crop-related data. Training such supervised models requires high quality labeled datasets. This work presents a dataset consisting of raw and preprocessed hyperspectral imaging (HSI) files capturing reflectance in the visible to near-infrared range (400-1000 nm) from two problematic weed species on California's Central Coast: annual sowthistle ( Sonchus oleraceus ) and little mallow ( Malva parviflora ). Hyperspectral imaging provides rich spectral-spatial data cubes that can support the development of deep learning models and autonomous technology for precision weed management. Plants were grown in a greenhouse under five conditions: standard, drought, overwatering, excess fertilizer, and no fertilizer. Custom MATLAB scripts were utilized for preprocessing, including k-means clustering to define regions of interest (ROIs), and extraction of spectral metrics. Data visualization was performed using Wolfram language and MATLAB. The dataset includes both raw and ENVI-formatted hyperspectral cubes and pre-processed MATLAB outputs, supporting spectral feature engineering, benchmark development, and exploratory machine learning workflows for controlled environment stress classification.

Why it matches plant phenotyping methods植物のストレス状態を対象とするハイパースペクトル画像データセットで、ROI抽出・スペクトル指標化と機械学習ベンチマークを中心的に提供しているため、植物フェノタイピング手法・データセットとして適格。

abstractThis work presents a dataset consisting of raw and preprocessed hyperspectral imaging (HSI) files capturing reflectance in the visible to near-infrared range (400-1000 nm) from two problematic weed species
Reproduction assets foundThe paper is a Data in Brief article describing its own hyperspectral imaging dataset of annual sowthistle and little mallow under five abiotic stress treatments, deposited publicly on Zenodo (record 17398082). The dataset includes raw ENVI-format hyperspectral cubes, preprocessed MATLAB outputs (ROI masks, extracted植被
Dataset · publicData accessibility Repository name: Zenodo Data identification number: zenodo.17398082 Direct URL to data: https://doi.org/10.5281/zenodo.17398082Open asset ↗Zenodo · zenodo.17398082html-lines:92-120
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 5 Sept 2026
Published15 May 2026Scientific ReportsCited by 0 · OpenAlex ↗

NucVerse3D: generalizable 3D nuclear instance segmentation across heterogeneous microscopy modalities.

MicroscopyX-ray / CTSegmentation

Accurate three-dimensional (3D) nuclear instance segmentation is a prerequisite for quantitative phenotyping in volumetric microscopy, yet remains challenging in densely packed tissues, irregular nuclear morphologies, and across heterogeneous imaging modalities. Here we present NucVerse3D, a deep-learning framework for generalized 3D nuclei instance segmentation that combines a residual attention 3D U-Net architecture with a reversible gradient-field representation for robust centroid-aware instance reconstruction. NucVerse3D is trained end to end in 3D using modality-agnostic preprocessing and isotropic scale normalization, enabling deployment across confocal microscopy, two-photon microscopy, light-sheet microscopy, micro-computed tomography, and scanning electron microscopy volumes. We benchmarked NucVerse3D on seven volumetric datasets spanning multiple species and tissues, comprising more than forty thousand manually annotated nuclei, including newly released ground-truth datasets of mouse liver tissue (control and hepatocellular carcinoma) and Drosophila brain glial nuclei. Across datasets, NucVerse3D achieved consistently high precision and competitive recall, resulting in strong F1-scores and average precision across a wide range of imaging conditions. While a modest precision-recall imbalance is observed in certain datasets, favoring high-confidence detections, this behavior reflects a conservative instance reconstruction strategy that prioritizes accurate boundary delineation and reduces false positive segmentation in densely packed and morphologically heterogeneous tissues. A single generalized model trained on pooled data matched the performance of dataset-specific models, and ablation experiments demonstrated that preprocessing and scale normalization substantially contribute to performance under strict intersection-over-union criteria. To demonstrate the biomedical utility of NucVerse3D, we applied it to 3D liver images from a mouse model of hepatocellular carcinoma (HCC) to enable spatially resolved 3D nuclear phenotyping. In healthy liver tissue, nuclear DNA content and nuclear volume exhibited a tightly regulated log-log scaling relationship. In contrast, tumor-adjacent and tumor regions displayed progressive disruption of this coupling, forming spatially coherent domains of nuclear DNA-volume decoupling that are not detectable in conventional two-dimensional histology. We quantify this phenomenon using a Nuclear Decoupling Score (NDS), revealing increased nuclear instability aligned with pathological tissue remodeling highlighting NDS as a potential quantitative biomarker of dysplastic and tumor tissue. Together, NucVerse3D provides a robust and generalizable solution for 3D nuclear instance segmentation and enables quantitative nuclear phenotyping across imaging modalities.

Why it matches plant phenotyping methods3D核セグメンテーション手法を開発し、多様な画像データセットでベンチマーク・検証したうえで、核形態とDNA量の定量的フェノタイピングに応用しており、植物対象ではないため本索引の対象外となる可能性はあるが、提示内容上はフェノタイピング手法研究として中心的である。

abstractAccurate three-dimensional (3D) nuclear instance segmentation is a prerequisite for quantitative phenotyping in volumetric microscopy
Reproduction assets foundThe paper's newly released Zenodo deposit (10.5281/zenodo.18517324) containing raw volumes, annotations, training patches, model weights, and segmentation outputs is not among the allowed URLs, so it cannot be listed. The authors' public analysis/segmentation code repository is explicitly deposited with an authors' URL
Code · publicThe source code for training and predicting nuclei segmentation using NucVerse 3D is available from https://github.com/Segovia-lab/3D-Nuclei-segmentation.git .Open asset ↗https://github.com/Segovia-lab/3D-Nuclei-segmentation.gitlines:647-728
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published15 May 2026PloS oneCited by 0 · OpenAlex ↗

A curated dataset and lightweight deep learning framework for tea leaf disease classification.

TeaLeafClassificationStress / disease detectionDisease symptoms / severity

Tea (Camellia sinensis) is the world's second most consumed beverage, enjoyed daily by more than two billion people. In Bangladesh, it serves as a cornerstone agricultural export and a major sector of the domestic economy. However, commercial tea cultivation remains highly vulnerable to fungal and pest-related diseases such as Blight, Red Rust, and Helopeltis which severely reduce crop yield and compromise leaf quality. While early detection is critical to preventing widespread outbreaks, traditional manual inspection is slow, subjective, and highly error-prone. Deep learning provides a scalable alternative, yet single-branch networks often struggle to capture both minute disease lesions and broader structural degradation simultaneously. To address this, we propose a Hybrid Feature Fusion architecture that runs two highly efficient feature extractors in parallel: EfficientNetV2-Small to isolate fine-grained local textures, and MobileNetV3-Small to capture the global structural context of the leaf. The models were trained and evaluated on a real-world dataset of 2,000 annotated images, evenly distributed across the four target classes (Blight, Red Rust, Helopeltis, and Healthy). Before training, the images underwent a standardized preprocessing pipeline including resizing to 224 × 224 pixels and normalization, supplemented by a dynamic augmentation strategy featuring random rotations, horizontal flips, and brightness adjustments to improve model robustness. The proposed hybrid framework achieved an outstanding peak classification accuracy of 96.80% alongside a macro Area Under the Curve (AUC) of 0.9980. To rigorously validate its performance, the hybrid model was benchmarked against six diverse architectures: a Vision Transformer (ViT-B16 at 76.40%), a Custom CNN (89.60%), MobileNetV3 (94.40%), ResNet50 (95.60%), DenseNet121 (96.40%), and EfficientNetV2-B3 (97.60%). Although EfficientNetV2-B3 achieved a marginally higher raw accuracy, the proposed dual-branch framework delivered a superior precision-recall balance and faster convergence stability. These findings demonstrate that the proposed hybrid methodology is highly reliable and computationally balanced, making it an ideal candidate for integration into Internet of Things (IoT) edge devices for real-time disease monitoring in precision agriculture.

Why it matches plant phenotyping methods茶葉の病徴を画像から分類する深層学習手法の開発と、注釈付きデータセットおよび複数モデルとのベンチマーク検証が中心であり、植物病害状態の表現型推定に該当する。

abstractwe propose a Hybrid Feature Fusion architecture
Reproduction assets foundThe paper's Data Availability statement explicitly deposits the curated 2000-image tea leaf dataset on Mendeley Data and the analysis code on GitHub, both with public URLs matching allowed_urls.
Dataset · publicThe dataset comprising 2000 annotated tea leaf images was curated under real-world field conditions. It has been made available at https://data.mendeley.com/datasets/3x42rbj8yv/1.Open asset ↗3x42rbj8yv/1html-lines:465-480
Code · publicThe computational code supporting the findings of this study is publicly accessible on GitHub: https://github.com/rayhankhan2192/Tea_Leaf_Disease_Model.Open asset ↗GitHub · rayhankhan2192/Tea_Leaf_Disease_Modelhtml-lines:465-480
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published8 May 2026Artificial Intelligence and ApplicationsCited by 0 · OpenAlex ↗

Classification of Multi-Crop Leaf Diseases in Rice, Wheat, and Bean Using a Deep Transfer Learning Approach

Common beanRiceWheatLeafClassificationDisease symptoms / severity

In Bangladesh, crop leaf diseases create a serious risk to food security and production from agriculture. Timely identification of leaf diseases in rice, wheat, and bean crops is considered crucial for the implementation of effective disease detection and classification strategies. To address this challenge, a MobilenetV2-based disease identification and classification system is proposed in this research. Previous studies focus on classifying diseases of a single species, leaving the need to train models separately for each species. This research focuses on forming a single standard model to perform leaf disease classification for multiple crop species including rice, wheat, and beans. The approach makes use of transfer learning with the MobilenetV2 model, which is fine-tuned using a dataset of annotated crop leaf images specific to Bangladesh. Following a comprehensive evaluation, an overall accuracy of 97.87% was achieved in the classification of crop leaf diseases, which surpasses the accuracy of a number of previous studies focusing on leaf disease detection of a single crop. The system demonstrates the capability to rapidly diagnose diseases in real time by enabling the users to prompt intervention to mitigate potential crop losses, ultimately leading to amplified crop yield and food security. Overall, the research highlights the promise of AI-powered solutions in tackling crop leaf disease detection, which in turn encourages greater research and technology adoption to support sustainable farming methods especially in the crop disease classification domain in Bangladesh and throughout the world. Received: 24 May 2025 | Revised: 9 March 2026 | Accepted: 14 April 2026 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement The data that support the findings of this study are openly available in the Bangladeshi Crops Disease Dataset at https://www.kaggle.com/datasets/nafishamoin/bangladeshi-crops-disease-dataset and the Bean Disease Dataset at https://www.kaggle.com/datasets/therealoise/bean-disease-dataset. Author Contribution Statement Md. Mahmudul Hasan: Conceptualization, Methodology, Visualization, Supervision. Md. Omar Faruq: Software, Validation, Writing – original draft. Mahadi Hasan Musa: Formal analysis, Investigation. Mohammad Mamunur Rashid: Resources, Data curation, Writing – review & editing. Khandaker Mohammad Mohi Uddin: Writing – review & editing, Project administration, Supervision.

Why it matches plant phenotyping methods葉画像から作物の病害状態を推定する深層学習手法を開発・評価しており、植物病害フェノタイピングが中心的な技術貢献である。

abstracta MobilenetV2-based disease identification and classification system is proposed in this research.
Reproduction assets foundThe paper's Data Availability Statement openly provides the Bean Disease Dataset on Kaggle, which is one of the two public image datasets used to train the multi-crop leaf disease classification model. The Bangladeshi Crops Disease Dataset URL is not among the allowed URLs, so only the bean dataset is reported. No code
Dataset · publict The authors declare that they have no conflicts of interest to this work. Data Availability Statement The data that support the findings of this study are openly available in the Bangladeshi Crops Disease Dataset at https:// www.kaggle.com/datasets/nafishamoin/bangladeshi-crops-disease- dataset and the Bean Disease Dataset at https://www.kaggle.com/datasets/therealoise/bean-disease-dataset.Author Contribution Statement Md. Mahmudul Hasan: Conceptualization, Methodology, Visualization, Supervision. Md. Omar Faruq: Software, Valida- tion, Writing – original draft. Mahadi Hasan Musa: Formal analysis, Investigation. Mohammad Mamunur Rashid: Resources, Data curation, Writing – review & editing.Open asset ↗Kaggle · therealoise/bean-disease-datasetpdf-raw-page:11 lines:1-83
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published7 May 2026PloS oneCited by 0 · OpenAlex ↗

Enhanced rice leaf disease classification via contour-driven segmentation and optimized deep transfer learning architectures.

RiceLeafClassificationStress / disease detectionDisease symptoms / severity

Pakistan is the fourth-largest rice producer and the fifth-largest exporter worldwide. Timely disease detection remains challenging due to the scale of cultivation and reliance on manual monitoring. Developing reliable, ongoing computerized systems for plant health management is essential for efficient disease control. A deep learning approach is used as the core method to identify diseases in rice leaves. This methodology employs a range of advanced deep learning architectures to achieve top-tier feature extraction and classification. The publicly available rice leaf disease dataset on Zenodo supports research reproducibility and data transparency. We systematically process a balanced dataset of 1914 image samples using Python with TensorFlow and a GPU to enable high-speed computation for large-scale image processing. This study conducts a systematic comparative evaluation of five deep transfer learning architectures (InceptionV3, DenseNet201, ResNet152V2, EfficientNetV2L and MobileNetV2) trained independently. The base backbone models are then integrated with guided GrabCut segmentation with contour-detection method for interpretable disease localization. In this work, the methods of segmentation by GrabCut and contour detection are introduced to make the results of the study easier to interpret and explain the disease areas, but the final classification outcomes are obtained only on the basis of the underlying deep transfer learning models. As a result, infected leaf areas can be identified more effectively, allowing for better understanding and explainable of the disease.To enhance interpretability, GrabCut segmentation and contour detection are applied as post-hoc visualization techniques to highlight diseased regions corresponding to CNN predictions. These techniques do not influence the classification training process. All five models InceptionV3, DenseNet201,ResNet152V2,EfficientNetV2L and MobileNetV2 demonstrated their effectiveness in detecting rice diseases during training, validation, and testing phases, with models trained over 30 epochs. The training methods and accuracy rates of the models were compared during validation and final testing. InceptionV3 demonstrated the most moderate performance of 98.80% training, 98.44% validation, and 98.43% test accuracy, which means that it has strong generalization and consistent learning behavior. The performance of very high-density networks such as DenseNet201 (98.72% train, 98.43% val, 98.43% test), ResNet152V2 (99.02% train, 99.22% val, 97.39% test), EfficientNetV2L model accuracies (39.01% train, 48.70% val, 44.50% test) also showed competitive results, which validated the effectiveness of deep transfer learning in the classification of rice leaf disease, while MobileNetV2 model accuracies (98.09% train, 98.18% val, 96.87% test) indicate that a lightweight model can still achieve reliable classification performance with lower computational complexity. In general, the comparative analysis defines InceptionV3 as the most stable and efficient model in the framework proposed. These results illustrate InceptionV3 superior generalization ability, supported by explainable methods for improved feature localization, confirming the viability of transfer learning for accurate and practical rice disease detection using GrabCut segmentation and contour detection technique. The complete implementation code and data used for the research experimentation is publicly available at https://github.com/ummershakeel03/Rice-Leaf-Diseases-Classification for reproducibility and reuse.

Why it matches plant phenotyping methodsイネ葉の病徴領域を画像から分類・局在化する深層学習ワークフローが研究の中心であり、GrabCut・輪郭検出と複数モデルの比較評価を含むため、植物病害状態の画像ベース表現型計測として採用。

abstractA deep learning approach is used as the core method to identify diseases in rice leaves.
Reproduction assets foundThe paper explicitly states that the complete implementation code and the rice leaf disease image dataset (1914 samples) used in this study are publicly available: code on the authors' GitHub repository and the dataset on Zenodo (DOI 10.5281/zenodo.15817084). Both are paper-specific, public, and actionable.
Code · publicThe complete implementation code and data used for the research experimentation is publicly available at https://github.com/ummershakeel03/Rice-Leaf-Diseases-Classification for reproducibility and reuse.Open asset ↗ummershakeel03/Rice-Leaf-Diseases-Classificationhtml-lines:1357-1368
Dataset · publicThe dataset for this research study is available at: https://doi.org/10.5281/zenodo.15817084.Open asset ↗10.5281/zenodo.15817084html-lines:1357-1368
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published6 May 2026Scientific dataCited by 1 · OpenAlex ↗

Morphometric Properties of Olive (Olea europaea) Pits: A Dataset for Cultivar Identification and Analysis.

OliveFruitClassificationMorphology / geometry measurementFruit / seed / panicle traits

Image analysis of pits and grains provide alternative routes for overcoming the invasive approach of genomic tools in the investigation of archaeological or modern plant material, which is only seldom a viable option due to the complex and laborious methodologies required. Nevertheless, any investigation of pit morphology and cultivar interpretation requires a high quality, comprehensive dataset for comparison. Such a benchmark dataset for the morphology of olive (Olea europaea) pits is presented in this paper, designed to facilitate similar research and establish a base for future investigations. The dataset was established by image analysis of pits of 18 olive cultivars that were photographed in both lateral and dorsal positions. A dedicated MATLAB® code was developed to extract the silhouettes of each pit and to calculate 16 morphometric traits of each view of the pit. Altogether, a total of 1008 photos of 504 pits of the 18 cultivars, together with their detailed morphometric description and statistical analysis are available here. These were used to test the accuracy of the dataset and the new approach in representing the different cultivars.

Why it matches plant phenotyping methodsオリーブ核の画像から形態形質を抽出する専用コードと、検証用ベンチマークデータセットを開発・提示しており、植物形質取得法が中心である。

abstractSuch a benchmark dataset for the morphology of olive (Olea europaea) pits is presented in this paper, designed to facilitate similar research and establish a base for future investigations.
Reproduction assets foundThe paper's olive pit images (1008 photos of 504 pits) and morphometric trait data (16 parameters per view) are openly deposited on Zenodo, along with the authors' MATLAB 'PitAnalyzer' software used for silhouette extraction and trait calculation. Both are paper-specific, public, and directly actionable via the Zenodo.
Dataset · publicAll the images are available on a dedicated Zenodo repository17. The file name of each image comprises an abbreviation of the cultivar name (Table 1), tree number (a, b or c), pit number (1–30) and the pit position (VD VL for dorsal and lateral, respectively).Open asset ↗Zenodohtml-lines:220-292
Code · publicThe code that was used in this work is compiled as a stand-alone software based on MATLAB “PitAnalyzer”. The software is available to download at the following repository, where any use of it should be attributed appropriately to this publication (https://zenodo.org/records/18789307).Open asset ↗Zenodohtml-lines:381-404
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 May 2026Scientific reportsCited by 1 · OpenAlex ↗

LDDHybridNet: an ROI-aware CNN-LSTM hybrid framework for accurate and early leaf disease detection in precision agriculture.

Field / plotLeafClassificationSegmentationStress / disease detectionDisease symptoms / severity

Early and accurate detection of plant leaf diseases is an essential requirement for precision agriculture, given their severe impact on global food security. While much has been done recently, many deep learning-based approaches will still fail in real-world tests because of challenges such as background clutter, differences in illumination, occlusion, or the fact that visual symptoms for these diseases can be very subtle early on. Traditional CNN- and Transformer-based architectures generally lack accurate lesion localisation and interpretability, hindering their practical deployment in agricultural decision-support tools. To address these issues, we present LDDHybridNet, a region-based, explanation-friendly deep learning framework that can identify leaf disease at an early, accurate stage. It then applies preprocessing steps guided by ROI, based on leaf segmentation from the U-Net, followed by a compact CNN-based spatial feature-extraction framework. We arrange spatial feature embeddings extracted from lesion regions into an ordered sequence and employ a Bi-LSTM with attention to model structured contextual dependencies, allowing progression-aware feature learning without requiring actual temporal image sequences. Lastly, Grad-CAM-based post-hoc explainability is employed to interpret model decisions, enabling transparent visualisation of disease-relevant regions. We conduct extensive experiments on the PlantVillage benchmark and the FieldPlant dataset and show that LDDHybridNet consistently outperforms representative CNN, transformer, and hybrid baselines across multiple evaluation metrics. Although the near-ceiling performance on PlantVillage reveals the dataset's artificial nature, the proposed framework achieves 95.37% accuracy under real-world field conditions and 92.84% on weak-lesion early-stage samples, demonstrating the method's robustness and early-stage detection potential. The performance boosts are statistically significant (P < 0.01). In general, LDDHybridNet is an interpretable and robust deep learning framework for leaf disease detection, which can support data-driven crop protection and precision agriculture applications.

Why it matches plant phenotyping methods葉の病害症状を画像から検出・局在化する深層学習手法の開発とベンチマーク評価が中心であり、植物の病害状態を直接推定するため、植物フェノタイピング手法として収録する。

abstractwe present LDDHybridNet, a region-based, explanation-friendly deep learning framework that can identify leaf disease at an early, accurate stage.
Reproduction assets foundThe paper's phenotyping measurements are leaf disease detection experiments on two public image datasets: PlantVillage (Kaggle) and FieldPlant (IEEE Dataport), both cited with explicit public URLs. The authors' code, trained weights, and scripts are not publicly released and are available only on request, so no code/模型
Dataset · public43.Hughes, D. P. & Mohanty, S. P. PlantVillage Dataset. [online] (2015). Available at: https://www.kaggle.com/datasets/emmarex/plantdiseaseOpen asset ↗PlantVillage Datasethtml-lines:657-726
Dataset · public44.Moupojou, R. K., Bouachir, W., Ahamed, T. & Taki, A. H. FieldPlant: A Real-World Dataset for Leaf Disease Detection in Field Conditions. IEEE Dataport. [online] (2021). Available at: https://ieee-dataport.org/documents/fieldplant-datasetOpen asset ↗FieldPlanthtml-lines:657-726
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published29 Apr 2026Frontiers in plant scienceCited by 0 · OpenAlex ↗

Comparative deep learning approaches for bean leaf disease recognition.

Common beanLeafClassificationStress / disease detectionDisease symptoms / severity

Context Plant diseases are a serious danger to the world's food security since they drastically lower crop output. Traditional manual plant leaf inspection is time-consuming, labor-intensive, and frequently subjective. Recent developments in deep learning provide effective and scalable methods for image-based analysis-based automated plant disease identification. Techniques Three deep learning architectures-a proprietary Convolutional Neural Network (CNN), ResNet18, and Vision Transformer (ViT)-are used in this study to examine automated bean leaf disease identification. The Augmented iBean dataset, which has three classes-angular leaf spot, bean rust, and healthy leaves-was used to train and assess the models. Every model was trained using the same preprocessing and training settings to provide fair benchmarking. Receiver Operating Characteristic (ROC) curves, accuracy, precision, and confusion matrices were used to assess the model's performance. Outcomes ResNet18 fared better than CNN and Vision Transformer models, according to a comparative analysis. ResNet18 maintained a high level of computing efficiency while achieving 99% accuracy and 99.01% precision. Its better categorisation capacity across all disease categories was validated using confusion matrix and ROC analysis. In conclusion The study shows that ResNet18 offers the optimal trade-off between accuracy and efficiency and creates a standard benchmarking framework for bean leaf disease identification. The results demonstrate its applicability for real-time deployment in precision agricultural systems for better crop management and early disease identification.

Why it matches plant phenotyping methods豆葉の病徴を画像から認識する深層学習手法を比較・ベンチマークしており、植物病害状態の取得手法が研究の中心である。

abstractThree deep learning architectures-a proprietary Convolutional Neural Network (CNN), ResNet18, and Vision Transformer (ViT)-are used in this study to examine automated bean leaf disease identification.
Reproduction assets foundThe paper's data availability statement points to the Augmented iBean dataset on IEEE DataPort, the public bean leaf image dataset used for all phenotyping/classification experiments in this study. No author analysis code or trained model checkpoints are shared.
Dataset · publicPublicly available datasets were analyzed in this study. This data can be found here: https://ieee-dataport.org/documents/bean-leaf-disease-augmented-ibean-dataset.Open asset ↗ieee-dataport · bean-leaf-disease-augmented-ibean-datasethtml-lines:446-496
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published28 Apr 2026Frontiers in plant scienceCited by 0 · OpenAlex ↗

Explainable deep learning-based comparative study for guava fruit and leaf disease classification: advancing agricultural diagnostics through AI.

FruitLeafClassificationStress / disease detectionDisease symptoms / severity

Introduction Early detection of plant diseases is essential for maintaining crop health and ensuring sustainable agricultural productivity. Guava fruit and leaf diseases, if not identified at an early stage, can lead to significant yield losses. Recent advances in deep learning offer promising solutions; however, challenges remain in achieving both high accuracy and model interpretability for practical agricultural deployment. Methods This study proposes an explainable deep learning-based framework for the classification of guava fruit and leaf diseases. A real-world dataset consisting of 527 annotated images across five classes-Disease Free, Phytophthora, Red Rust, Scab, and Styler and Root Rot-was utilized. Six hybrid model architectures were developed by integrating transfer learning backbones (VGG16, MobileNetV2, InceptionV3, and ResNet50) with custom convolutional neural network (CNN) classifiers. Model performance was evaluated using accuracy, precision, recall, F1-score, and class-wise metrics. To enhance transparency, Gradient-weighted Class Activation Mapping (Grad-CAM) was employed to visualize disease-relevant regions. Results Among all evaluated models, the proposed VGG16 + MobileNetV2 hybrid architecture achieved the best performance, attaining an accuracy of 96%, an F1-score of 0.96, and strong generalization across all disease classes. Comparative analyses using confusion matrices, ROC-AUC curves, precision-recall curves, and radar plots confirmed the superior and consistent performance of the proposed model over other hybrid configurations. Discussion The results demonstrate that combining deep feature extractors with lightweight architectures enhances both classification accuracy and computational efficiency. The integration of Grad-CAM provides meaningful visual explanations, increasing trust and interpretability in AI-assisted disease diagnosis. This framework shows strong potential for deployment in real-time smart farming systems and mobile-based diagnostic applications, particularly in resource-constrained agricultural environments.

Why it matches plant phenotyping methodsグアバの葉・果実画像から病害状態を推定する深層学習手法が研究の中心であり、複数モデルの比較評価とGrad-CAMによる説明可能性検証も行っているため、植物フェノタイピング方法論として採用する。

abstractThis study proposes an explainable deep learning-based framework for the classification of guava fruit and leaf diseases.
Reproduction assets foundThe paper's plant image dataset (527 annotated guava fruit/leaf disease images) is a public Kaggle deposit explicitly cited by the authors with a URL, making it a paper-specific, publicly actionable asset. No author analysis code or trained model checkpoints are stated as publicly available; the data availability only指
Dataset · publicKaggle ). Available online at: https://www.kaggle.com/datasets/noamaanabdulazeem/guava-dataset (Accessed January 10, 2024 ).Open asset ↗Kagglelines:550-617
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published24 Apr 2026PloS oneCited by 0 · OpenAlex ↗

PlantaNet and PlantaNetLite: Efficient and explainable multi-crop plant disease classification via transformer benchmarking and custom lightweight CNNs.

Whole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

Plant disease diagnosis based on visual symptoms is crucial for preventing yield loss; however, deployment in practical settings remains challenging due to inter-class similarity, background noise, and limited computational resources. This study presents a plant disease classification framework evaluated on a curated multi-crop dataset aggregated from multiple publicly available repositories, comprising 51 disease and healthy classes. The dataset includes approximately 45,000 original images that were expanded through controlled augmentation during training to improve generalization. We benchmark eight ImageNet-pretrained tiny vision transformer architectures trained for up to 50 epochs. Among these, CAFormer-s18 achieved strong validation performance but with increased computational overhead. To enable efficient and computationally lightweight solutions, we design two fully customized convolutional neural networks: PlantaNetLite (1.28M parameters) and PlantaNet (2.58M parameters). After hyperparameter optimization and full 100-epoch training, PlantaNet achieved 99.37% validation accuracy and 99.66% test accuracy with a compact model size (9.85 MB) and moderate computational cost, while PlantaNetLite achieved a best validation accuracy of 99.22% under further parameter reduction. Qualitative Grad-CAM and Grad-CAM++ analyses provide insight into the regions influencing model predictions. Overall, the proposed models demonstrate competitive accuracy while maintaining computational efficiency, highlighting their potential suitability for resource-constrained deployment scenarios.

Why it matches plant phenotyping methods植物の視覚症状から病害状態を推定する画像ベースの表現型解析手法を開発・比較し、データセット上で性能評価しているため、方法が中心的である。

abstractThis study presents a plant disease classification framework evaluated on a curated multi-crop dataset aggregated from multiple publicly available repositories
Reproduction assets foundThe paper's Data Availability Statement explicitly states the curated multi-crop plant disease image dataset used for all classification experiments is publicly available on Kaggle at the authors' URL. No author analysis code, trained model checkpoints, or code repository is disclosed in the supplied blocks.
Dataset · publicThe dataset used in this study is publicly available at https://www.kaggle.com/datasets/alimransonet/plant-disease-dataset.Open asset ↗Kaggle · alimransonet/plant-disease-datasethtml-lines:727-758
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 5 Sept 2026
Published23 Apr 2026Plant phenomics (Washington, D.C.)Cited by 0 · OpenAlex ↗

MSDDG: Multi-scale dual-discriminator GAN for point cloud completion of plant

Eggplant / auberginePumpkin / squashSunflowerLiDAR / point cloudWhole plant / canopy / plot / field2D/3D reconstructionArchitecture / morphology / geometry

Plant 3D reconstruction using optical imaging often suffers from incomplete point clouds due to viewpoint occlusion and sensor limitations. This incompleteness hinders accurate structural representation and subsequent feature extraction for plant analysis. To address these challenges, we propose a Multi-Scale Dual-Discriminator Generative Adversarial Network (MSDDG) for plant point cloud completion. A multi-scale point cloud generator (MSPG) that integrates local and global features from raw incomplete point clouds is used for MSDDG to reconstruct complete shapes. The dual-discriminators-a multi-view projected silhouette discriminator and a spatial distance discriminator-are designed to ensure geometric realism and spatial plausibility from multiple perspectives. To train MSDDG, we created the Plant4L dataset containing four plant species (sunflower, pumpkin, luffa, and eggplant) with high-quality 3D models augmented via 3D thin plate spline transformations and virtual occlusion simulation to generate incomplete point clouds and multi-view silhouettes. Experimental results on Plant4L demonstrate that MSDDG achieves superior completion performance, with Chamfer Distance (CD), Hausdorff Distance (HD), and Uniformity Chamfer Distance (UCD) all below 0.41. Comparative evaluations confirm MSDDG's superiority over previous point cloud completion methods. The application of MSDDG for 3D reconstruction from single view further validate its effectiveness in restoring occluded plant structures.

Why it matches plant phenotyping methods植物の不完全点群を補完して3D構造を再構成する手法を開発し、植物データセット上で比較評価・検証しており、表現型取得ワークフローが中心です。

abstractwe propose a Multi-Scale Dual-Discriminator Generative Adversarial Network (MSDDG) for plant point cloud completion.
Reproduction assets foundThe paper's data availability statement explicitly states that the source code and datasets (including the Plant4L point cloud completion dataset) are publicly available at the authors' GitHub repository.
Code · publicThe source code and datasets used in this study are publicly available at https://github.com/Amuro-Aznable/MSCGPCN.git .Open asset ↗https://github.com/Amuro-Aznable/MSCGPCN.git · MSCGPCNlines:415-421
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 5 Sept 2026
Published20 Apr 2026Plant MethodsCited by 1 · OpenAlex ↗

A systematic comparison of transformers and ConvNets for root segmentation across nine datasets.

RootMorphology / geometry measurementSegmentationRoot system architecture

BACKGROUND: Root segmentation is a fundamental yet challenging task in image-based plant phenotyping. Accurate segmentation is a prerequisite for extracting root traits relevant to plant physiology, breeding, and agronomy. While U-Net and other convolutional neural network (ConvNet) architectures have been applied to root segmentation, no systematic comparison of multiple Transformer and ConvNet architectures has been conducted across diverse root imaging conditions. RESULTS: We evaluated 21 segmentation architectures across nine diverse root image datasets, training 1511 models to assess all combinations of architecture, dataset, pre-training strategy, and learning rate, producing over 3 million segmentations for evaluation. Transformer-based models significantly outperformed ConvNets for Dice (mean Dice 0.679 vs 0.659; [Formula: see text]). Root-diameter and root-length correlation were also higher for Transformers, but the differences were not statistically significant ([Formula: see text] and [Formula: see text] respectively). Pre-training significantly improved mean Dice from 0.623 to 0.666 ([Formula: see text]), with Transformers benefiting more from pre-training than ConvNets (Dice improvement + 0.072 vs + 0.021; [Formula: see text]), supporting the hypothesis that fine-tuned Transformers transfer more effectively across large domain gaps. MobileSAM achieved the highest Dice score (0.693) while maintaining computational efficiency. Both architecture families underestimated thin root length compared to manual annotations. Dataset choice explained 70.9% of performance variance, far exceeding model architecture (6.7%). PURPOSE: Transformer architectures significantly outperform ConvNets for root segmentation accuracy, and pre-training significantly improves performance, particularly for Transformers. Pre-trained MobileSAM offers the best accuracy at competitive computational cost. Dataset choice dominates performance variance, suggesting practitioners should prioritize data curation over architecture selection.

Why it matches plant phenotyping methods根の画像セグメンテーション手法を複数データセットで体系的に比較・検証し、根長・根径などの形質抽出性能も評価しているため、植物フェノタイピング手法が中心である。

abstractRoot segmentation is a fundamental yet challenging task in image-based plant phenotyping.
Reproduction assets foundThe paper's root image datasets (DeepRootLab, Grassland, Chicory, PRMI) are publicly available, and the authors' training code and modified RhizoVision Explorer trait-extraction fork are on GitHub with explicit availability statements.
Dataset · publicImages are available from https://zenodo.org/records/15213661 .Open asset ↗Zenodo · 15213661lines:872-982
Dataset · publicImages are available from https://figshare.com/ndownloader/articles/20440497/versions/2 .Open asset ↗Figshare · 20440497lines:872-982
Dataset · publicImages are available from https://zenodo.org/records/3527713 .Open asset ↗Zenodo · 3527713lines:872-982
Dataset · publicImages are available from https://gatorsense.github.io/PRMI/ .Open asset ↗lines:872-982
Code · publicTraining code is available at https://github.com/sotlampr/seg .Open asset ↗GitHub · sotlampr/seglines:1183-1225
Code · publicAll nine root image datasets used in this study are publicly available. DeepRootLab images are available from Zenodo (https://zenodo.org/records/15213661). Grassland images are available from Figshare (https://figshare.com/ndownloader/articles/20440497/versions/2). Chicory images are available from Zenodo (https://zenodo.org/records/3527713). The six PRMI datasets (Papaya, Peanut, Sesame, Sunflower, Cotton, Switchgrass) are available from https://gatorsense.github.io/PRMI/. Training code is available at https://github.com/sotlampr/seg. The modified RhizoVision Explorer fork used for trait extraction is available at https://github.com/sotlampr/RhizoVisionExplorer.Open asset ↗GitHub · sotlampr/RhizoVisionExplorerlines:1294-1347
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 Apr 2026Plant methodsCited by 0 · OpenAlex ↗

Deep learning-based identification of visually similar foliar diseases in field-grown barley.

BarleyField / plotLeafSegmentationStress / disease detectionDisease symptoms / severity

Background Accurate segmentation of foliar diseases under field conditions is essential for large-scale phenotyping, as breeding programs rely on reliable severity estimates to identify genotypes with improved resistance. However, most deep learning approaches have been developed as pathogen-specific models, which limits scalability in field-grown barley where multiple diseases naturally co-occur and exhibit substantial visual similarity. Results We evaluated whether a multiclass segmentation model can simultaneously detect and distinguish two fungal diseases of barley, Puccinia hordei and Ramularia collo-cygni, and compared its performance with two disease-specific binary models. Using 336 high-resolution leaf scans collected in the field with naturally occurring co-infections, the multiclass model achieved higher Dice scores for brown rust (0.59 vs 0.40; +47.5% relative improvement) and ramularia (0.60 vs 0.53; +13.2% relative improvement). It also captured a greater proportion of individual lesions across both classes. At the genotype level, the model-predicted disease area percentages were highly consistent with those from ground truth annotations ([Formula: see text]). Conclusions A unified multiclass framework can more effectively segment visually similar foliar diseases than separate binary models, while simplifying the computational workflow. This provides a scalable basis for automated resistance assessment within breeding pipelines. Code and data are publicly available at https://github.com/grimmlab/BarleyDiseaseSegmentation, with Mendeley Data dataset DOI 10.17632/4ny92p2r8f.1.

Why it matches plant phenotyping methods圃場画像から葉面病害面積をセグメンテーションし、遺伝子型レベルの病害重症度を推定する手法を開発・比較・検証しており、植物フェノタイピングが中心です。

abstractAccurate segmentation of foliar diseases under field conditions is essential for large-scale phenotyping
Reproduction assets foundThe paper's annotated barley leaf disease segmentation dataset (Mendeley Data DOI 10.17632/4ny92p2r8f.1) and the authors' analysis/segmentation code (GitHub grimmlab/BarleyDiseaseSegmentation) are explicitly declared publicly available, directly reproducing this paper's phenotyping measurements and computational models
Dataset · publicThe annotated dataset and the code implementing our machine learning–based model are publicly available on Mendeley Data (https://doi.org/10.17632/4ny92p2r8f.1) and GitHub (https://github.com/grimmlab/BarleyDiseaseSegmentation).Open asset ↗Mendeley Data · 10.17632/4ny92p2r8f.1lines:133-140
Code · publicCode and data are publicly available at https://github.com/grimmlab/BarleyDiseaseSegmentation, with Mendeley Data dataset DOI 10.17632/4ny92p2r8f.1.Open asset ↗GitHub · grimmlab/BarleyDiseaseSegmentationlines:1-70
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published17 Apr 2026Plant phenomics (Washington, D.C.)Cited by 2 · OpenAlex ↗

BloomSight: An ultra-high-frequency phenotyping framework for diurnal flowering dynamics in japonica and indica rice to enable genetic dissection and hybrid-breeding applications.

RiceFlowerPanicle / ear / spikeMorphology / geometry measurementObject detectionSegmentationGrowth / development / phenologyFruit / seed / panicle traits

Rice ( Oryza sativa ) production underpins food security in many rice-consuming nations. As a critical developmental transition that directly determines yield and grain quality, flowering dates and timing are genetically complex and highly sensitive to environmental fluctuations. This complexity requires new methods to quantify diurnal floral characteristics, which are essential to hybrid breeding in cereals. Here, we present BloomSight, an ultra-high-frequency and deep-learning (DL) powered framework for phenotyping and measuring minute-level flowering dynamics in japonica and indica rice. After monitoring 172 rice accessions selected from the Chinese Rice Mini-Core Collection using cost-effective time-lapse imaging platforms for 16 days, we acquired over 530,000 accession-level images and established the Open Rice Flowering Training (ORFT) dataset, with over 39,000 panicles and 350,000 anthers annotated. Next, a two-stage customised DL model (i.e. YOLACT-Panicle for panicle segmentation and UNet-Anther for anther identification) was trained using the ORFT set, enabling ultra-high-frequency measures of anther extrusion at the minute level. Based on trait analysis, we further fitted curves to dynamically identify diurnal flowering patterns, including key timepoints such as the initial flowering timepoint ( T Ini. ), quickest flowering timepoint ( T Qck. ), and peak flowering time ( T Peak ), and novel traits such as the duration of rapid flowering phase ( P Rpd. ) and flowering density across key phases. After validating BloomSight-derived traits against manual observations, we classified the japonica and indica accessions into three patterns: Slow, Moderate, and Fast, all of which had distinct flowering windows. These analyses helped us integrate phenotypic variations into a genome-wide association study (GWAS), revealing many significant single nucleotide polymorphisms (SNPs) associated with known (e.g. EMF1 , OsMYB8 , and PME42 ) and several repeatedly identified unknown loci (one of these loci has been recently verified by other groups), demonstrating the value of the BloomSight framework. Taken together, we believe that BloomSight provides an ultra-high-frequency framework for diurnal flowering phenotyping, enabling the measurement of biological meaningful floral traits with minute-level resolution that can enable flowering-related developmental studies and hybrid-breeding applications in rice and more broadly benefit the plant and crop research community.

Why it matches plant phenotyping methodsイネの開花動態を高頻度画像と深層学習で抽出するフェノタイピング基盤を開発し、データセット構築と手動観測による検証も行っているため、方法が研究の中心である。

abstractwe present BloomSight, an ultra-high-frequency and deep-learning (DL) powered framework for phenotyping and measuring minute-level flowering dynamics in japonica and indica rice
Reproduction assets foundThe paper's Data and code availability statement explicitly provides public access to the ORFT annotated image dataset (BioStudies S-BSST2157), Python source code for floral trait analysis (GitHub The-Zhou-Lab/BloomSight), and trained DL models (GitHub releases). SRA accessions are molecular sequencing data, not phenot
Code · publicPython-based source codes for automating floral trait analysis using the above data are accessible via our GitHub repository ( https://github.com/The-Zhou-Lab/BloomSight ).Open asset ↗The-Zhou-Lab/BloomSightlines:336-349
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published13 Apr 2026Data in briefCited by 1 · OpenAlex ↗

BDFlower: Growth stage flower image dataset for precision agriculture and floriculture.

RGB / grayscaleFlowerClassificationGrowth / development / phenology

This study presents a comprehensive BDFlower growth stage dataset designed to support research in precision agriculture and floriculture. The dataset encompasses eight common flower species found in Bangladesh: Bush Allamanda, Red Hibiscus, Yellow Bell, Pinwheel Flower, Pink Periwinkle, White Madagascar Periwinkle, Marvel of Peru, and White Hibiscus. Each species is represented across three growth stages-Early, Mid, and Full-resulting in 24 distinct classes. A total of 23,334 colour images are included, comprising 3889 original photographs and 19,445 augmented samples generated with five augmentation techniques. Bush Allamanda contains 499 images, Red Hibiscus contains 489 images, Yellow Bell contains 483 images, Pinwheel Flower contains 497 images, Pink Periwinkle contains 452 images, White Madagascar Periwinkle contains 472 images, Marvel of Peru contains 468 images and White Hibiscus contains 529 images. Each image was collected using smartphone camera at three-time intervals per day, spaced eight hours apart, to capture natural variations in lighting and appearance. The dataset is further organized into training, validation, and testing splits, enabling direct application to machine learning workflows. This is a publicly available dataset specifically curated for flower growth stage classification. In addition to dataset collection, we also conducted a simple experiment using a CNN model to evaluate its performance on this dataset. It is intended to facilitate the development of robust computer vision models that can monitor flower development, with potential applications in automated plant phenotyping, crop monitoring, and digital floriculture systems.

Why it matches plant phenotyping methods花の生育段階を画像で分類する公開データセットを構築し、CNN評価も行っており、植物表現型取得・解析が研究の中心である。

abstractThis is a publicly available dataset specifically curated for flower growth stage classification.
Reproduction assets foundThe paper's own flower growth-stage image dataset (BDFlower) is publicly deposited on Mendeley Data with an explicit direct URL and DOI, directly reproducing the paper's phenotyping (flower growth stage) image measurements. No author analysis code or trained model checkpoints are explicitly deposited.
Dataset · publicRepository name: Data Mendeley Data identification number: 10.17632/m8g2wynwyr.2 Direct URL to data: https://data.mendeley.com/datasets/m8g2wynwyr/2Open asset ↗10.17632/m8g2wynwyr.2html-lines:94-129
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 Apr 2026Data in briefCited by 0 · OpenAlex ↗

A benchmark dataset of Primitive Indian Paddy Panicle Images and identification via deep residual transfer learning.

RicePanicle / ear / spikeClassificationFruit / seed / panicle traits

We introduce ``Primitive Indian Paddy Panicle Images,'' a benchmark image dataset of 22 primitive Indian rice panicle varieties (Sethy, Prabira; Pamerelli, Ranjith, 2026; Mendeley Data, V1, doi:10.17632/khfd7pzskd.1) and present an identification approach based on deep residual transfer learning. Using a transfer-learned ResNet-50 with image augmentation and an 80/10/10 train/validation/test split, the model attains 100.0% validation accuracy and 98.74% accuracy on the held-out test set. Per-class one-vs-rest AUCs on validation are 1.000 for all 22 classes; test AUCs range from 0.9924 to 1.000 (mean ≈ 0.999), with separate confusion matrices and ROC curves provided for validation and test partitions. These results demonstrate that deep residual transfer learning can robustly discriminate closely related panicle morphotypes when trained on a carefully curated dataset. We release the dataset to support reproducible research in germplasm identification, varietal purity assessment, and automated phenotyping.

Why it matches plant phenotyping methodsイネ穂画像のベンチマークデータセットと、深層学習による穂形態の自動識別手法が研究の中心であり、再現可能な植物表現型解析基盤として明示されている。

abstractWe introduce ``Primitive Indian Paddy Panicle Images,'' a benchmark image dataset of 22 primitive Indian rice panicle varieties
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicDirect URL to data: https://data.mendeley.com/datasets/khfd7pzskd/1Open asset ↗Mendeleyhtml-lines:1-116
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published12 Apr 2026Plant, cell & environmentCited by 0 · OpenAlex ↗

Plant Species With an Acquisitive Resource-Use Strategy Exhibit Lower Wood Density and Display Greater Intraspecific Variation.

LeafStem / branchPhysiological trait estimationLeaf traitsWater status / transpiration

Leaf and hydraulic traits are key determinants of growth rates, and hence potentially exhibit significant associations with wood density (WD) and its intraspecific variation (ITV). However, the extent to which functional traits could improve WD prediction accuracy, and how ITV in WD correlates with functional traits remain incompletely understood. We investigated WD and its ITV across 10,218 plant species, mapped the global distribution of WD, and analyzed the association of ITV in WD with niche breadth and functional traits. Plant species with an acquisitive resource-use strategy, characterized by higher specific leaf area (SLA), leaf nitrogen concentration (LN), and leaf maximum stomatal conductance (g max ), exhibited lower WD. Associations of WD with hydraulic traits indicated species with greater hydraulic safety exhibited higher WD. Moreover, the integration of leaf traits (i.e., SLA and LN) and hydraulic traits with environmental factors substantially enhanced WD prediction accuracy in a random forest model, raising the explained variance from 55% to 95%. Furthermore, resource-acquisitive species demonstrated higher ITV for WD. ITV was positively related to relative niche breadth concerning both climatic factors and soil properties. Overall, functional traits significantly improve WD prediction accuracy, and plant species with an acquisitive resource-use strategy exhibit lower WD but greater intraspecific variation.

Why it matches plant phenotyping methods木材密度という植物形質の予測モデルを構築し、機能形質・環境因子の統合による予測精度を検証しており、形質推定手法が中心的です。

abstractthe integration of leaf traits (i.e., SLA and LN) and hydraulic traits with environmental factors substantially enhanced WD prediction accuracy in a random forest model, raising the explained variance from 55% to 95%.
Reproduction assets foundThe paper's Data Availability Statement points to a public Zenodo deposit containing the authors' global wood density distribution data, which directly reproduces this paper's measurements. The TRY Plant Trait Database is a generic third-party database, not a paper-specific asset, and no author analysis code is stated.
Dataset · publicData for the global distribution of wood density is available on Zenodo Repository https://sandbox.zenodo.org/records/425279.Open asset ↗Zenodo · 425279html-lines:405-429
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published11 Apr 2026Scientific reportsCited by 0 · OpenAlex ↗

Frost damage segmentation in grapevine organs using YOLOv11s with ASPP and dynamic confidence thresholding.

GrapevineField / plotFruitLeafSegmentationStress / disease detectionStress response / tolerance

Climate change, particularly increasing frequency and intensity of spring frost events, poses a serious threat to viticulture by reducing yield and product quality. This study proposes an image processing and machine learning-based framework for early, rapid, and accurate segmentation of frost damage in vineyards using YOLOv11s enhanced with Atrous Spatial Pyramid Pooling (ASPP). A unique dataset called FGVL dataset from Sultana seedless grape vineyards in Manisa, Türkiye, following a severe frost event in April 2025. FGVL includes 418 frost-damaged grapes, 510 frost-damaged leaves, 395 healthy grapes, and 698 healthy leaves, all manually annotated by experts under natural field conditions. By integrating ASPP into YOLOv11s, proposed model improved multi-scale contextual feature extraction and achieved mAP@50 of 0.7686, demonstrating stronger performance in instance segmentation of small, overlapping, and visually similar grapevine organs. In addition, Dynamic Confidence Thresholding (DCT) strategy was introduced to improve prediction reliability in dense and visually complex vineyard scenes. Despite challenges such as background clutter, object overlap, and small target structures, model maintained stable performance with low computational demand, requiring only 6.45 GB of GPU memory. Proposed framework offers an accurate, efficient, and practically deployable early recognition system for frost damage assessment in viticulture.

Why it matches plant phenotyping methodsブドウの器官における霜害状態を画像からセグメンテーションする手法を開発・評価しており、植物の病害・障害状態の取得が研究の中心である。

abstractThis study proposes an image processing and machine learning-based framework for early, rapid, and accurate segmentation of frost damage in vineyards using YOLOv11s enhanced with Atrous Spatial Pyramid Pooling (ASPP).
Reproduction assets foundThe paper's Data availability statement explicitly shares the FGVL frost-damage dataset and source code in the corresponding author's public GitHub repository, matching an allowed URL.
Code · publicSource code and dataset are publicly shared in GitHub repository of corresponding author. GitHub repo: https://github.com/kaanarikk/Grape-Instance-Segmentation-For-ViticultureOpen asset ↗https://github.com/kaanarikk/Grape-Instance-Segmentation-For-Viticulturelines:230-236
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 Apr 2026Data in briefCited by 0 · OpenAlex ↗

Dataset of RGB images of healthy grapevine leaves and with downy mildew, powdery mildew, Esca complex, and erineum mite symptoms.

GrapevineField / plotRGB / grayscaleLeafClassificationDisease symptoms / severity

This dataset consists of a collection of high-resolution RGB images of grapevine leaves, designed to support research in plant pathology, precision viticulture, and computer vision. The images were collected in situ from experimental and commercial vineyards in the north of Portugal, covering different vineyard conditions and management practices. The dataset includes healthy leaves from three grapevine Portuguese cultivars Loureiro, Viosinho and Malvasia Fina, photographed under natural lighting conditions without artificial adjustments. It is organized into five categories: healthy leaves and leaves showing symptoms of downy mildew ( Plasmopara viticola ), powdery mildew ( Erysiphe necator ), Esca complex and Erineum Mite ( Colomerus vitis ). Images are provided in JPEG format with a resolution of 3000 × 3000 pixels and 1024 × 1024 pixels and arranged in folders by health status and disease type. This dataset can be used for machine learning and deep learning applications in disease detection/classification, cultivar identification, and can support other precision agriculture applications, as well as being used for agricultural robotics and educational purposes. An evaluation on three deep learning architectures demonstrated the suitability of the dataset into separating the five classes.

Why it matches plant phenotyping methodsブドウ葉の病徴を画像化した再利用可能なデータセットで、植物の健康状態・病害状態の画像ベース推定を支えることが中心です。深層学習による5クラス分類評価も記載されています。

abstractThis dataset consists of a collection of high-resolution RGB images of grapevine leaves, designed to support research in plant pathology, precision viticulture, and computer vision.
Reproduction assets foundThe paper is a Data in Brief article describing a public Zenodo repository of RGB grapevine leaf images (healthy plus downy mildew, powdery mildew, Esca complex, erineum mite) collected for plant disease/phenotyping research, with explicit data accessibility details. No author analysis code or trained model checkpoints
Dataset · publicData accessibility Repository name: Zenodo Data identification number: https://doi.org/10.5281/zenodo.17343473 Direct URL to data: https://zenodo.org/records/17343473Open asset ↗Zenodo · 10.5281/zenodo.17343473html-lines:93-144
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published30 Mar 2026Cited by 0 · OpenAlex ↗

A Comprehensive Image Dataset of Fruit and Leaf Diseases Across Six Horticultural Crops for Deep Learning Applications

AppleBanana / plantainCitrusMangoRGB / grayscaleFruitLeafClassificationStress / disease detectionDisease symptoms / severity

Abstract Abstract. Accurate and timely identi cation of plant diseases is essential for improving crop productivity and ensuring sustainable agricultural practices. This paper presents a comprehensive image dataset of fruit and leaf diseases covering six economically important horticultural crops: Apple, Banana, Citrus, Guava, Mango, and Papaya. The dataset comprises high-quality RGB images representing both healthy and diseased samples, with disease symptoms including spots, lesions, discoloration, blight, rot, and fungal and bacterial infections captured under diverse real-world conditions. Variations in illumination, background complexity, viewing angles, growth stages, and symptom severity are intentionally included to enhance the robustness and generalizability of learning models developed using this data. The dataset is structured in a class-wise manner and preprocessed to support direct integration with deep learning frameworks. It is extensively used to train, validate, and evaluate deep learning based plant disease classi cation models, enabling automatic feature learning from raw images without manual intervention. Experimental usage demonstrates that the dataset is well suited for convolutional neural networks and attentionbased architectures, facilitating e ective discrimination between multiple disease categories across di erent crops and plant organs. By providing a uni ed multi-crop, multi-disease benchmark, this dataset aims to accelerate research in automated crop disease diagnosis, precision agriculture, and intelligent decision-support systems for sustainable farming.

Why it matches plant phenotyping methods植物の葉・果実の病徴画像を収録したデータセット/ベンチマークであり、病害状態の画像ベース推定を中心的に扱うため。

abstractThis paper presents a comprehensive image dataset of fruit and leaf diseases covering six economically important horticultural crops
Reproduction assets foundThe paper's core asset is the ABCGMP fruit and leaf disease image dataset, publicly deposited on Mendeley Data, with author analysis code also stated to be available on GitHub. Both are paper-specific, public, and actionable.
Dataset · publicData is available on Mendeley:1Open asset ↗pdf-page:33 lines:1-56
Code · publicCode availability: Code is available on GitHub 2Open asset ↗GitHubpdf-page:33 lines:1-56
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published27 Mar 2026Plants (Basel, Switzerland)Cited by 1 · OpenAlex ↗

TB-DLossNet: Fine-Grained Segmentation of Tea Leaf Diseases Based on Semantic-Visual Fusion.

Field / plotMultimodalLeafSegmentationDisease symptoms / severity

Camellia oleifera is an economically vital woody oil crop. Its productivity and oil quality are severely compromised by various diseases. Implementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection. Despite recent progress, existing segmentation methods struggle with three primary challenges: semantic ambiguity arising from evolving pathological stages, blurred boundaries due to overlapping lesions, and the high omission rate of micro-lesions. To address these issues, this paper presents TB-DLossNet (Text-Conditioned Boundary-Aware Network with Dynamic Loss Reweighting), a novel segmentation framework based on semantic-visual multi-modal fusion. Leveraging VMamba as the visual backbone, the proposed model innovatively integrates BERT-encoded structured text as an auxiliary modality to resolve visual ambiguities through cross-modal semantic guidance. Furthermore, a boundary enhancement branch is incorporated alongside a multi-scale deep supervision strategy to mitigate boundary displacement and ensure the topological continuity of lesion structures. To tackle the detection of small-scale targets, we designed a dynamic weight loss function conditioned on lesion area, significantly bolstering the model's sensitivity to minute pathological features. Additionally, to alleviate the scarcity of high-quality data, we curated a comprehensive multi-modal dataset encompassing seven typical diseases of Camellia oleifera . Experimental results demonstrate that TB-DLossNet achieves a Mean Intersection over Union (mIoU) of 87.02%, outperforming the state-of-the-art unimodal VMamba and multimodal Lvit by 4.9% and 2.59%, respectively. Qualitative evaluations confirm that our model exhibits lower false-negative rates and superior boundary-fitting precision in heterogeneous field scenarios. Finally, generalization tests on an apple disease dataset further validate the robustness and transferability of the proposed framework.

Why it matches plant phenotyping methods植物病害の病斑を画素レベルで抽出する新規セグメンテーション手法を開発し、データセット整備と性能比較・汎化検証も行っているため、病害状態の画像ベース表現型計測が中心である。

abstractImplementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection.
Reproduction assets foundThe authors state their code and experimental dataset (the multimodal Camellia oleifera disease segmentation dataset) are publicly available on GitHub, matching an allowed URL.
Code · publicOur code and experimental dataset are available at https://github.com/zzzsq239/TB-1.Open asset ↗zzzsq239/TB-1html-lines:820-841
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published25 Mar 2026Frontiers in plant scienceCited by 2 · OpenAlex ↗

YOLOv11n-DualPC-Lite: a lightweight, high-precision real-time detection model for maize leaf diseases.

MaizeField / plotLeafObject detectionDisease symptoms / severity

To address the challenge of balancing model lightweight and detection accuracy in maize leaf disease detection, as well as the limitations of edge device deployment resources, we propose an enhanced target detection model, YOLOv11n-DualPC-Lite.Firstly, the C2fDualPConv module was designed, integrating PartialConv to replace some C3k2 modules in the backbone and neck networks. This approach enhances feature representation while reducing the number of parameters. Secondly, the Slim-Neck architecture is introduced in the neck network. To improve accuracy without increasing the number of parameters, the VoVGSCSPC_SimAm module enables the new Slim-Neck structure to reduce parameters while strengthening feature representation. Finally, an EfficientHead detection head is introduced that uses an inverted bottleneck MBConv module to improve performance. This significantly reduces computational load while efficiently extracting features. This study constructed a maize leaf disease dataset integrating a publicly available Kaggle dataset and a field-collected dataset from Anhui Science and Technology University's experimental plots. The dataset includes four categories: Blight, Common_Rust, Gray_Leaf_Spot, and Health. Through techniques such as rotation and gamma correction, the dataset was expanded from 3,876 to 5,165 images for model training and performance validation. Test results show this improved model performs better than other popular lightweight models overall, with a mAP50 score of 90.9%. Meanwhile, the model has only 2.13 million parameters; its computational complexity is reduced to 4.55 G, and the model size is 4.41 MB. Compared with the original YOLOv11n, its mAP50 is 1.9% higher, while the number of parameters is down by 17.8%, computational complexity is cut by 29.3%, and file size is reduced by 15.7%. When run on a Raspberry Pi 5, the model's detection speed reaches 2.3 FPS, an increase of 27.8%. This model achieves a good balance between detection accuracy and lightweight performance for maize leaf diseases, providing an efficient and practical method for real-time crop disease monitoring.

Why it matches plant phenotyping methodsトウモロコシ葉の病徴を画像から検出・分類する軽量モデルを開発し、データセット構築、性能比較、エッジデバイス検証まで行っており、植物病害状態の画像ベース表現型取得が中心である。

abstractwe propose an enhanced target detection model, YOLOv11n-DualPC-Lite
Reproduction assets foundThe paper's maize leaf disease detection study uses a public Kaggle maize leaf disease image dataset (Dataset 1) combined with a field-collected dataset. The Kaggle dataset is a public, paper-specific image asset directly used for the model's training and validation. No author analysis code, trained model checkpoints,或
Dataset · publicre, the model was successfully run on a Raspberry Pi 5 edge device, realizing stable, real-time detection and providing a workable technical method for field disease monitoring. 2 Materials and methods 2.1 Dataset introduction The dataset constructed in this study comprises two datasets: Dataset 1 from the Kaggle data website ( https://www.kaggle.com/datasets/hendriyunuswijaya/maize-leaf-disease ) and Dataset 2 collected from the experimental field at Anhui Science and Technology University in Chuzhou City, Anhui Province. Dataset 1 contains a total of 4,188 images, including 1,162 images in the Health category. All images depict only specific regions of healthy maize leaves without complex Open asset ↗Kaggle · hendriyunuswijaya/maize-leaf-diseaselines:46-63
Code / dataset availability confirmedEurope PMC · bioRxiv · checked 5 Sept 2026
Published23 Mar 2026bioRxivCited by 0 · OpenAlex ↗

Quantification of anatomical changes in young grapevine wood over time and in response to Neofusicoccum parvum with image processing

GrapevineMicroscopyTissueMorphology / geometry measurementArchitecture / morphology / geometry

Grapevine Trunk diseases (GTDs) represent a major threat for the wine industry. Despite several break-through, their etiology remains unclear and no curative treatment is currently available. Wood anatomy and water transport contribute to the symptoms of young plant decline. This study investigates wood anatomical alterations in two Alsatian grapevine cultivars presenting different susceptibility to GTDs, focusing on wood structure over six months of vegetative growth and in response to infection. Using a validated FasGa staining protocol, wood sections from transverse, tangential, and radial directions were stained to differentiate lignified and cellulosic tissues. Microscopic analysis was performed at x4, x10, and x40 magnifications, yielding a dataset of 4771 images. To support this high-throughput quantitative analysis of microscopy images, a computational model was developed, enabling reliable and efficient assessment of anatomical traits. Pre-established woody tissues presented higher xylem vessels diameter in Gewurztraminer than Riesling, with a dorsoventral arrangement whereas the number of vessels remained the same all over the cross section. No significant anatomical changes were observed in established woody tissues, whereas newly formed xylem anatomy showed a possible rearrangement during infection, especially in Gewurztraminer cultivar. Furthermore, colorimetric analysis quantified the lignification of woody tissues in response to wounding damage compared to un-treated plants. While definitive conclusions remain limited due to the experimental timeframe and sample variability, the findings highlight the need for longer-term studies and broader cultivar evaluation. Code and microscopy images have been made publicly available, providing a scalable digital tool for future research in plant vascular systems.

Why it matches plant phenotyping methods植物組織画像から木部解剖形質と木化を定量する計算モデルを開発・検証し、大規模画像データセットと公開コードを提供しており、表現型取得手法が研究の中心である。

abstractTo support this high-throughput quantitative analysis of microscopy images, a computational model was developed, enabling reliable and efficient assessment of anatomical traits.
Reproduction assets foundThe paper's microscopy image dataset (4771 grapevine wood images) is publicly deposited on Zenodo with an explicit DOI matching an allowed URL. The authors also state their Python analysis pipeline is available at github.com/courbot/vineside, but that URL is not among the allowed URLs, so only the Zenodo image dataset,
Dataset · publicThis database can benefit the research community, and is publicly available online at https://doi.org/10.5281/zenodo.18850060 [35].Open asset ↗Zenodo · 10.5281/zenodo.18850060pdf-page:4 lines:1-56
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 Mar 2026Frontiers in plant scienceCited by 0 · OpenAlex ↗

AutoSiQ: a curated haploid Arabidopsis thaliana inflorescence dataset with a fine-grained silique ontology and a deep learning application for haploid fertility quantification.

ArabidopsisFlowerFruitPanicle / ear / spikeClassificationCountingObject detectionFruit / seed / panicle traits

Doubled haploid (DH) technology can fast-track crop breeding. Haploid induction yields haploids with only one set of genomes, which are usually sterile. Haploid fertility (HF) is the ability of haploid plants to set seed, and it is a critical bottleneck in DH pipelines. Genetic mechanisms to restore HF hold immense potential in DH crop breeding, yet its phenotyping remains manual, destructive, and inconsistent. While recent advances in imaging and machine learning have improved throughput for general plant traits, no curated image dataset exists for Arabidopsis thaliana that explicitly represents HF. Here, we present AutoSiQ, a dataset and baseline deep learning pipeline for automated HF quantification. AutoSiQ includes high-resolution scanned inflorescences annotated with a seven-class ontology encompassing green siliques, green fertile siliques, mature siliques, fertile siliques, cracked fertile siliques, cracked siliques, and flowers. This multi-class annotation scheme preserves biologically meaningful information beyond binary fertile/non-fertile distinctions, enabling reliable fertility estimation and future phenotyping applications. We release baseline object detection models (YOLOv5), trained using the AutoSiQ dataset, and evaluate their performance across confidence thresholds. Model predictions strongly correlate with manual counts, achieving R² up to 0.94 for total silique number estimation. We further demonstrate AutoSiQ's utility for automated haploid fertility rate (HFR) estimation and genotype discrimination between two contrasting genotypes (WT and bmf2 mutant). A longitudinal analysis identifies ~60 days after sowing (DAS) as the optimal harvest time for maximizing mature silique counts by balancing between the number of immature buds and silique shattering. By releasing both the dataset and baseline code, AutoSiQ provides a reproducible and extensible foundation for high-throughput fertility phenotyping in haploid Arabidopsis .

Why it matches plant phenotyping methodsハプロイド稔性を画像から定量するデータセットと深層学習パイプラインを開発・評価しており、植物フェノタイピング手法が中心である。

abstractHere, we present AutoSiQ, a dataset and baseline deep learning pipeline for automated HF quantification.
Reproduction assets foundThe paper's AutoSiQ dataset (annotated scanned Arabidopsis inflorescence images with seven-class silique ontology and manual fertility counts) is publicly deposited on Zenodo per the data availability statement. The YOLOv5 GitHub repository is a generic third-party library, not an authors' code asset.
Dataset · publicThe datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: https://zenodo.org/records/17905566 .Open asset ↗zenodo · 17905566lines:367-402
Code / dataset availability confirmedEurope PMC · bioRxiv · Crossref · checked 14 Sept 2026
Published18 Mar 2026bioRxivCited by 0 · OpenAlex ↗

Spectral Phenotyping Reveals Time-Specific QTLs in Field-Grown Lettuce

LettuceField / plotMultispectral / hyperspectralWhole plant / canopy / plot / fieldPhysiological trait estimationGrowth / time-series analysisGrowth / development / phenologyStress response / tolerance

Lettuce ( Lactuca sativa ) is an important field crop, but our understanding of its phenotypic variation and underlying genetics under natural field conditions remains limited, posing challenges for identifying effective crop breeding targets. Longitudinal hyperspectral phenotyping allows for non-invasive monitoring of crop performance under diverse agricultural conditions. In this study, we used hyperspectral imaging to assess the phenotypic variation of almost 200 different field-grown lettuce varieties, following the same plants from just after seedling- to flowering-stage. With automated image processing, we extracted a wide range of spectral phenotypes related to metabolite content, growth efficiency, and environmental stress responses, creating a multi-dimensional time-resolved data set. Principal component analysis (PCA) revealed the major axes of spectral variation over time, and highlighted differences in spectral patterns among lettuce genotypes. Integrating on-site weather data, we modelled G×E interactions of reflectance, revealing regions of the lettuce vegetation spectrum that are primarily shaped by genotype and/or environment. We estimated phenotypic plasticity in response to time, temperature and rainfall using best linear unbiased predictions (BLUPs), capturing genotype-specific developmental trajectories and responses to the environment. We used genome-wide association studies (GWAS) to identify quantitative trait loci (QTLs) of PC-based, single and BLUP-based phenotypes, disentangling the genetic architecture of spectral lettuce phenotypes from major axes of variation down to single wavelength spectral plasticity. These findings provide new insights into the genome-wide genetic regulation and dynamics of spectral phenotypes in field grown lettuce.

Why it matches plant phenotyping methods圃場レタスを対象に、縦断ハイパースペクトル画像と自動画像処理でスペクトル形質を抽出するフェノタイピング手法・データセットが研究の中心である。

abstractLongitudinal hyperspectral phenotyping allows for non-invasive monitoring of crop performance under diverse agricultural conditions.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicScripts used in this study can be found at https://github.com/SnoekLab/Hyperspec_Mehrem_etal_2025.Open asset ↗SnoekLab/Hyperspec_Mehrem_etal_2025pdf-page:9 lines:1-31
Code / dataset availability confirmedOpenAlex · Crossref · checked 13 Sept 2026
Published18 Mar 2026Springer Science and Business Media LLCCited by 0 · OpenAlex ↗

Hierarchically scaled remote sensing and field datasets for three-dimensional wildland fuel characterization

Aerial / UAVField / plotPhotogrammetry / SfM / MVSLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurementCalibration / preprocessing2D/3D reconstructionArchitecture / morphology / geometry

Abstract Background: Next-generation models of fire behavior and smoke production rely on gridded, 3D inputs of wildland fuel complexes. We used a hierarchically scaled sampling design to characterize canopy and surface fuels that are common to prescribed burning programs in the southeastern and western US. Sampling included airborne laser scanning, terrestrial laser scanning, close-range photogrammetry, and destructive field sampling. The objective of this study was to use a combination of airborne laser scanning (ALS), terrestrial laser scanning (TLS), structure-from-motion photogrammetry (SfM), and field observations to create co-located 3D datasets of live and dead understory fuels for use in wildland fuel mapping and prescribed burn decision support Results: Using our integrated, co-located methods, we produced hierarchically-scaled datasets detailing the structure and composition of canopy and surface fuels across 9 southeastern pine sites, 5 western pine sites, and 4 western grassland sites. These are now publicly available at within the Wildland Fire Science Initiative data repository (https://doi.org/10.60594/W4859C). In this paper, we detail methods and the repository structure. Conclusions: The study was designed to evaluate and advance methods for 3D fuel characterization and to provide consistently scaled and labelled datasets for model training and evaluation. More specifically, machine learning models can be used to parse 3D point clouds collected from ALS, TLS, and structure-from-motion photogrammetry into fuel objects and metrics. Calibration with field plots will allow our hierarchically-scaled datasets to be used as the foundation for synthetic fuelbed mapping, starting with fine-scale objects such as individual shrubs or downed wood and scaling to vegetation patches and operational burn units.

Why it matches plant phenotyping methodsALS、TLS、SfMと現地観測を統合して植物群落の3D構造・燃料特性を取得し、手法の評価・改良と公開データセット構築を主目的としているため、植物形質計測法が中心である。

abstractThe objective of this study was to use a combination of airborne laser scanning (ALS), terrestrial laser scanning (TLS), structure-from-motion photogrammetry (SfM), and field observations to create co-located 3D datasets of live and dead understory fuels for use in wildland fuel mapping and prescribed burn decision support
Reproduction assets foundThe paper's hierarchically scaled ALS/TLS/SfM point clouds, field fuel measurements, and analysis scripts are explicitly stated to be open source and archived in the Wildland Fire Science Initiative data repository (DOI 10.60594/W4859C), a paper-specific public asset directly reproducing this study's phenotyping/fuel-3
Dataset · publicThe datasets and analysis scripts for this study are open source and are being archived with the Wildland Fire Science Initiative data repository (doi.org/10.60594/W4859C), including project metadata, methods documentation and data libraries (Prichard and Rowell 2025).Open asset ↗Wildland Fire Science Initiative data repository · 10.60594/W4859Clines:384-403
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published13 Mar 2026Cited by 0 · OpenAlex ↗

A Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis

Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationGrowth / development / phenology

Abstract Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.

Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画のデータセットで、植物フェノロジー自動検出の訓練、検証、ベンチマークを目的とする方法論的成果である。

titleA Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis
Reproduction assets foundThe paper is a Data Note describing the Multi-Modal Actinidia chinensis Phenology Dataset, which is explicitly stated to be publicly available on Zenodo with a DOI matching an allowed URL. The dataset contains the paper's own phenotyping assets: 1,665 annotated images with bounding-box phenological labels, georeferened
Dataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available at Zenodo: https://doi.org/10.5281/zenodo.17371025. This dataset comprises two components: (1) 1 665 JPEG images (1 024 × 1 024 pixels) with corresponding Pascal VOC XML annotation files containing bounding box coordinates and phenological class labels, and (2) 24 MP4 video files (3 840 × 2 160 pixels) with corresponding GPX coordinate files and Excel validation files containing manual ground truth counts.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:13 lines:1-62
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published9 Mar 2026Data in briefCited by 0 · OpenAlex ↗

Vegetation dynamics inside Mediterranean vineyards: A dataset for tracking changes using unmanned aerial vehicles.

GrapevineAerial / UAVField / plotPhotogrammetry / SfM / MVSLiDAR / point cloudRGB / grayscaleMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldClassification

Service crops are grown to provide ecosystem services in viticulture, but their adoption remains limited due to their competition with grapevine for soil resources. To identify trade-offs between services, the effect of service crops management strategies on grapevine performances still need further research. This dataset presents data from two experiments conducted to study the effect of service crops management on soil resources and grapevine performances. The inter-row vegetation was sampled in two Mediterranean vineyards using quadrats for biomass estimation. In addition, an unmanned aerial vehicle (UAV) was regularly flown over the vineyards for a period spanning more than four years in total over the two vineyards. The dataset presented here includes both raw data acquired during fieldwork and processed data derived from this raw inputs. The raw data consists of image series captured by two UAVs during each flight campaign, including RGB and multispectral imagery. Images were acquired between 2021-06-10 and 2022-07-29 for the first vineyard, and between 2023-06-08 and 2025-03-12 for the second vineyard. Based on these raw data, the processed data comprises spatial vectors, raster layers, and dense point clouds generated from UAV images using a Structure from Motion (SfM) photogrammetry workflow, at a 5 cm spatial resolution. The raster layers and dense point clouds provide specific information on vineyard characteristics for each UAV flight date, including elevation, vegetation indices, visible and near-infrared reflectance, and canopy height. In addition, the processed data include measurements of vegetation dry biomass, as well as separate measurements of dry biomass and leaf area measured for selected service crops species. This dataset can be reused for the calibration and/or evaluation of classification algorithms aimed at discriminating vines from the inter-row vegetation, or as part of a larger dataset to explore relationships between remotely-sensed vegetation indices and field-measured vegetation biomass or surface.

Why it matches plant phenotyping methodsUAV画像とSfM処理により、植生指数・樹冠高・バイオマス等の植物形質を取得した再利用可能なデータセットで、分類アルゴリズムの校正・評価用途も明示されており、植物フェノタイピング手法・データ基盤が中心です。

abstractThe dataset presented here includes both raw data acquired during fieldwork and processed data derived from this raw inputs.
Reproduction assets foundThe paper is a Data in Brief article describing a public dataset on Research Data Gouv (doi: 10.57745/MXM55R) containing UAV RGB/multispectral imagery, SfM-derived rasters and point clouds, and field-measured vegetation biomass/leaf-area data from two Mediterranean vineyards — directly the paper's phenotyping inputs. A
Dataset · publicollected in vineyards located in southern France near Montpellier (43°32.5243′N, 3°50.8240′E). Data are stored on Research Data Gouv, a remote storage solution curated by the French Department of Research. Data accessibility Repository name: Research Data Gouv Data identification number: doi: 10.57745/MXM55R Direct URL to data: https://doi.org/10.57745/MXM55R Related research article None 1. Value of the Data • The fine scale imaging of vineyards (i.e., 5 cm resolution) allows for classification of the vegetation in the vineyard inter-rows, and subsequent exploration of its respective dynamics. •Open asset ↗Research Data Gouv · 10.57745/MXM55Rlines:1-47
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published6 Mar 2026Plant phenomics (Washington, D.C.)Cited by 1 · OpenAlex ↗

Quantification of lettuce leaf DUS test traits and phenotypic fingerprint construction for variety identification.

LettuceLeafClassificationMorphology / geometry measurementSegmentationLeaf traitsPigment / colour / senescence

Rapid and accurate identification of DUS (Distinctness, Uniformity, and Stability) test traits in lettuce leaves is essential for advancing multi-omics-driven intelligent breeding. It also plays a critical role in germplasm protection and enhancing agricultural competitiveness. However, the phenotypic traits of lettuce leaves are highly diverse and complex due to both genotypic variation and environmental influences, posing significant challenges for precise DUS trait quantification. To address these challenges, we propose a high-precision phenotypic trait extraction pipeline and introduce an interpretable phenotypic fingerprinting framework for lettuce subgroup identification. First, a lightweight semantic segmentation network guided by group attention is developed to extract leaf components. Then, shape, color, and texture traits are comprehensively quantified. Following UPOV (International Union for the Protection of New Varieties of Plants) guidelines, we establish quantitative methods for seven DUS test traits: leaf shape, leaf tip shape, leaf margin shape, leaf vein shape, color hue, brightness, and anthocyanin coloration. Finally, PCA (Principal component analysis) was used to select 13 key traits, capturing over 95.82% of the total variance, for constructing "phenotypic ID" of lettuce varieties. Experiments conducted on 709 lettuce leaf image datasets showed that the accuracy of subgroup identification based on phenotypic fingerprints reached 98.59%. This study offers a scalable approach for automated DUS test trait evaluation and intelligent crop variety identification, providing a novel paradigm with strong potential for application in precision breeding and germplasm resource management.

Why it matches plant phenotyping methodsレタス葉画像からDUS形質を抽出・定量化する画像解析パイプラインを開発し、709画像で評価しており、植物フェノタイピング手法が研究の中心である。

abstractwe propose a high-precision phenotypic trait extraction pipeline and introduce an interpretable phenotypic fingerprinting framework for lettuce subgroup identification.
Reproduction assets foundThe article provides a public GitHub repository containing the authors' source code for the lettuce phenotypic fingerprint pipeline. The 709-image dataset and annotations are only available upon request, so they do not qualify as public assets.
Code · publicThe data used to support the findings of this study are available upon request from the corresponding author, and the source code is accessible at https://github.com/qiuguangjie87/PP_Phenotypic_Fingerprint .Open asset ↗PP_Phenotypic_Fingerprintlines:263-278
Code / dataset availability confirmedOpenAlex · arXiv · checked 15 Sept 2026
Published4 Mar 2026arXiv (Cornell University)Cited by 0 · OpenAlex ↗

CLIP-Guided Multi-Task Regression for Multi-View Plant Phenotyping

LeafWhole plant / canopy / plot / fieldMorphology / geometry measurementGrowth / development / phenologyLeaf traits

Modeling plant growth dynamics plays a central role in modern agricultural research. However, learning robust predictors from multi-view plant imagery remains challenging due to strong viewpoint redundancy and viewpoint-dependent appearance changes. We propose a level-aware vision language framework that jointly predicts plant age and leaf count using a single multi-task model built on CLIP embeddings. Our method aggregates rotational views into angle-invariant representations and conditions visual features on lightweight text priors encoding viewpoint level for stable prediction under incomplete or unordered inputs. On the GroMo25 benchmark, our approach reduces mean age MAE from 7.74 to 3.91 and mean leaf-count MAE from 5.52 to 3.08 compared to the GroMo baseline, corresponding to improvements of 49.5% and 44.2%, respectively. The unified formulation simplifies the pipeline by replacing the conventional dual-model setup while improving robustness to missing views. The models and code is available at: https://github.com/SimonWarmers/CLIP-MVP

Why it matches plant phenotyping methods植物画像から葉数・植物齢を推定するマルチビュー表現学習手法を開発し、ベンチマークで性能評価しており、表現型取得・推定が研究の中心である。

abstractWe propose a level-aware vision language framework that jointly predicts plant age and leaf count using a single multi-task model built on CLIP embeddings.
Reproduction assets foundThe paper explicitly states that the model and code are publicly available at the authors' GitHub repository, which qualifies as a paper-specific public code asset.
Code · publicm 7.74 to 3.91 and mean leaf-count MAE from 5.52 to 3.08 compared to the GroMo baseline, corresponding to improvements of 49.5% and 44.2%, respectively. The unified formulation simplifies the pipeline by replacing the conventional dual-model setup while improving robustness to missing views. The modela and code is available at: https://github.com/SimonWarmers/CLIP-MVP Index Terms: Plant phenotyping, Multi-view learning, Multi-task regression, Precision agriculture † † address: † Computer Vision Lab, CAIDAS, IFI, University of Würzburg, Germany ‡ Technological University Dublin, Ireland 1 Introduction Plant phenotyping from multiview imagery is crucial for precision agriculture, enabling non-Open asset ↗SimonWarmers/CLIP-MVPlines:1-53
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published3 Mar 2026Plant phenomics (Washington, D.C.)Cited by 1 · OpenAlex ↗

Low-annotation apple flower counting: A color-SAM enhanced and uncertainty-guided semi-supervised framework.

AppleAerial / UAVRGB / grayscaleFlowerCountingSegmentationFruit / seed / panicle traits

Accurate flower-load assessment is critical for informed thinning strategies in orchard management. UAV-based deep learning automated counting offers efficiency advantages, yet precise counting is heavily dependent on abundant annotated data, which is scarce and costly to obtain in agricultural settings. While semi-supervised learning alleviates dependency on manual annotation, its application to UAV-based orchard imagery faces challenges: complex backgrounds and small target sizes, which undermine pseudo-label reliability. To address these challenges, this study proposes a two-stage framework to achieve separate counting of apple flowers at different phenological stages. First, a color-SAM flower extractor (CSAM-FE) is proposed to preprocess images using a strategy combining color thresholding with the Segment Anything Model (SAM), suppressing background noise and extracting high-quality flower clusters, thereby providing purified inputs for the subsequent counting network. Second, an uncertainty-guided semi-supervised flower counting network (USCount-Net) is proposed for accurate stage-specific flower counting with limited labeled data. The USCount-Net incorporates two key components: an adaptive pseudo-label filtering (PLF) mechanism based on frequent forward uncertainty estimation (FFUE) is designed to dynamically suppress noisy gradient backpropagation, mitigating error propagation from unreliable pseudo-labels; and a noise-sensitive adaptive gated fusion (AGF) module is introduced to fuse cross-scale features without redundancy, addressing significant scale variations across phenological stages and observation angles. Comparative experiments on a self-built apple flower counting dataset demonstrate that USCount-Net achieves lower MAE and RMSE than state-of-the-art methods at 10%, 30%, and 50% labeling ratios. The results demonstrate that the proposed methodology serves as methodological support for rapid and precise apple flower counting in low-annotation agricultural scenarios.

Why it matches plant phenotyping methodsリンゴ花の画像抽出・計数手法と半教師あり解析ネットワークを開発し、データセット上で比較評価しているため、植物表現型取得が中心である。

abstractthis study proposes a two-stage framework to achieve separate counting of apple flowers at different phenological stages.
Reproduction assets foundThe paper's Data availability statement explicitly provides public access to the authors' USCount-Net source code on GitHub and the self-built apple flower counting dataset (UAV images, annotations, flower cluster images) on Google Drive.
Code · publicThe source code is publicly available at https://github.com/haohuihui5019/USCount-Net . And the source dataset can be accessed at https://drive.google.com/drive/folders/1KP8H0qIuct56hWre5GV6ZJnzwOpen asset ↗USCount-Netlines:681-780
Dataset · publicThe source code is publicly available at https://github.com/haohuihui5019/USCount-Net . And the source dataset can be accessed at https://drive.google.com/drive/folders/1KP8H0qIuct56hWre5GV6ZJnzwOpen asset ↗lines:681-780
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 Mar 2026Data in briefCited by 0 · OpenAlex ↗

A LiDAR-based machine vision dataset for online volume measurement of sweetpotatoes.

LiDAR / point cloudRGB / grayscaleRootMorphology / geometry measurementSegmentationArchitecture / morphology / geometry

Volume is an important shape descriptor in postharvest quality evaluation and breeding programs of sweetpotatoes and is also valuable for other agricultural engineering applications. Traditional volume measurement methods based on water displacement are, however, laborious, destructive, and unsuitable for high-throughput online scenarios. To address this gap, this dataset was developed to support the advancement of non-destructive, automated online volume estimation using a LiDAR (light detection and ranging)-based three-dimensional (3-D) machine vision system. A total of 200 sweetpotato storage roots of the cultivar "Beauregard" were collected for constructing a 3-D multi-view imagery dataset. Each sample was imaged online using a short-range LiDAR camera (Intel RealSense™ L515) while traveling on a custom-built roller conveyor system that enables simultaneous translation and rotation for full-surface coverage. The curated dataset comprises raw color images (1280 × 720 pixels, .png format) and corresponding raw and segmented point clouds (1280 × 720 pixels, .laz format) for individual samples, alongside the reference volume measurements obtained using the standard water displacement method. In addition, to illustrate the modeling pipeline for volume prediction, the dataset provides the extracted geometric features derived from the segmented two-dimensional (2-D) masks and point clouds, and volume prediction results obtained through regression modeling. As the first publicly available LiDAR-based dataset for sweetpotato volume estimation, this dataset provides a valuable resource for developing and validating image processing pipelines, optimizing machine learning models, and advancing 3-D vision technologies for non-destructive, rapid measurement of the volume of irregularly shaped agricultural products.

Why it matches plant phenotyping methodsサツマイモ貯蔵根の体積という植物器官形質をLiDAR 3D画像から推定する公開データセットであり、取得系・参照測定・特徴抽出・予測結果を含むため、フェノタイピング手法とデータセットが中心です。

abstractthis dataset was developed to support the advancement of non-destructive, automated online volume estimation using a LiDAR (light detection and ranging)-based three-dimensional (3-D) machine vision system.
Reproduction assets foundThe paper's own LiDAR sweetpotato dataset (images, point clouds, ground-truth volumes, feature data, and Python modeling scripts) is publicly deposited on Zenodo with an explicit DOI. The librealsense GitHub link is a generic camera SDK, not a paper-specific asset.
Dataset · publicDirect URL to data: https://doi.org/10.5281/zenodo.18378019Open asset ↗Zenodo · 10.5281/zenodo.18378019html-lines:90-113
Code · publicThe complete Python modeling script and the associated feature datasets have been included in the public dataset repository [13] to facilitate reproducibility and provide a benchmark for future algorithm development.Open asset ↗html-lines:168-182
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 Mar 2026Plant phenomics (Washington, D.C.)Cited by 1 · OpenAlex ↗

Synthetic-augmented multimodal deep learning fuses dual-angle RGB images and phenology to unlock genotype-informative canopy structural trait in wheat.

WheatField / plotMultimodalRGB / grayscaleWhole plant / canopy / plot / fieldMorphology / geometry measurementGrowth / time-series analysisArchitecture / morphology / geometryGrowth / development / phenologyYield / yield components

The wheat canopy genome harbors abundant yet untapped genetic variation that could be harnessed to enhance yield potential. The green area index (GAI) is a structural metric that reflects the photosynthetically active canopy surface and is closely linked to final grain yield. Current image-based GAI retrieval methods often suffer from signal saturation and coarse structural depiction, constraining downstream genetic analyses. To address this limitation, we constructed a comprehensive image dataset spanning eight field experiments across China and France, encompassing approximately 600 genotypes under six distinct management regimes. Leveraging this diverse data, we developed a multimodal deep-learning framework augmented by simulated-to-realistic (sim2real) synthetic data transfer. This framework fuses nadir and oblique RGB images with accumulated thermal time to produce high-precision, time-series GAI estimates. Validated on independent testing datasets from both China and France, the multimodal approach demonstrated robust performance with an accuracy of R 2 = 0.88 and an RMSE of 0.49 m 2 m -2 , representing an improvement of about 22% over the traditional gap fraction method. In three site-year field experiments involving 565 genotypes, the GAI dynamics derived from the multimodal approach showed higher broad-sense heritability (0.20-0.48) than those from the gap fraction approach (0.02-0.13) and stronger genotypic correlations with yield (0.19-0.40 versus 0.09-0.31). Furthermore, genetic analysis confirmed the biological fidelity of the estimated traits, identifying loci that co-localize with known architectural regulators such as Rht-D1 , TaTB1-4D , and TaBGC1-4D . Consistently, the multimodal-derived phenotypes were specifically enriched in cell-wall remodeling and hormonal signaling pathways (e.g., brassinosteroid) that directly regulate canopy expansion. Overall, the proposed method offers a powerful tool for unlocking genetic gain in canopy architecture and accelerating canopy-targeted wheat improvement.

Why it matches plant phenotyping methodsデュアルアングルRGB画像と熱時間を統合してGAIを推定する深層学習法を開発し、独立データで検証しているため、植物形質取得法が研究の中心です。

abstractwe constructed a comprehensive image dataset spanning eight field experiments across China and France
Reproduction assets foundThe paper publicly releases its pre-trained multimodal GAI-estimation model weights and inference code on Hugging Face, directly reproducing this paper's phenotyping analysis. The raw image and phenology datasets are not public and require contacting the authors.
Code · publicThe pre-trained model weights, inference code, and usage instructions are publicly available in the Hugging Face repository at https://huggingface.co/PheniX-Lab/GAI-Estimation/tree/main .Open asset ↗PheniX-Lab/GAI-Estimationlines:259-277
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published2 Mar 2026Data in briefCited by 0 · OpenAlex ↗

Seasonal collection of in situ optical and thermal images dataset and meteorological measurements over an Indian semi-arid rice crop.

RiceField / plotMultimodalMultispectral / hyperspectralThermalWhole plant / canopy / plot / fieldCalibration / preprocessingLeaf traitsPlant / canopy heightPlant / canopy temperature

This article describes a multi-sensor dataset collected during the TIRAMISU (Thermal InfraRed Anisotropy Measurements in India and Southern eUrope) campaign at the Nawagam research site in Gujarat, India, during the 2023 monsoon season. The objective was to acquire continuous ground-based optical and thermal measurements over a homogeneous rice canopy across different crop growth stages. The dataset integrates several complementary components. Thermal data were acquired with an Optris longwave infrared camera (8-14 µm) at high temporal resolution, capturing canopy temperature dynamics throughout the diurnal cycle. Optical data were obtained with a Micasense RedEdge-M multispectral sensor, providing imagery in Blue, Green, Red, RedEdge, and Near-Infrared bands with radiometric corrections. An Apogee radiometer supplied reference radiometric temperature. Meteorological measurements included air temperature, humidity, wind speed and direction, and net radiation. Ancillary field measurements comprised Leaf Area Index (LAI), plant height, emissivity sampling, hyperspectral observations, and crop stage information. The datasets are provided with metadata and processing workflows, including calibration procedures for optical reflectance and thermal radiance. Together, these components form a comprehensive record of canopy-atmosphere interactions over a homogeneous rice field. The datasets can support research on optical and thermal directional anisotropy, canopy radiative transfer, emissivity characterization, and crop biophysical parameter estimation. In addition, they are relevant for applications in vegetation monitoring, agricultural water stress assessment, and surface energy balance studies. By combining optical, thermal, and meteorological observations, the resource is suited for multidisciplinary investigations in remote sensing, agronomy, and environmental sciences.

Why it matches plant phenotyping methods光学・熱画像、校正手順、処理ワークフロー、LAIや草丈などの植物形質を含む再利用可能な作物キャノピーデータセットが研究の中心であり、植物表現型取得基盤として適格。

abstractThe dataset integrates several complementary components.
Reproduction assets foundThe paper is a Data in Brief describing the TIRAMISU rice-canopy dataset (thermal/multispectral images, meteorological, ancillary LAI/height, hyperspectral, emissivity) publicly deposited at doi.org/10.6096/1028, including processing scripts (Thermal_CSV_to_Image.py, MicaSense notebook) for reproducibility.
Dataset · publicRepository name: Optical, Thermal Infrared, and Meteorological Dataset from the Thermal InfraRed Anisotropy Measurements in India and Southern eUrope (TIRAMISU) Rice Canopy Experiment Data identification number: doi.org/10.6096/1028 Direct URL to data: https://doi.org/10.6096/1028 Instructions for access: Publicly accessible repository; representative subsets provided with metadata and processing scripts. Related research article Pinnepalli, C., Roujean, J.-L., Irvine, M., et al. [ 1 ]. Measuring and modelling directional effects in the frame of TIRAMISU. ISPRS Annals, X–3–2024 , 325–330. https://doi.orgOpen asset ↗doi.org · 10.6096/1028lines:49-77
Code / dataset availability confirmedOpenAlex · Crossref · Europe PMC · checked 5 Sept 2026
Published28 Feb 2026Scientific DataCited by 0 · OpenAlex ↗

Tomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping.

TomatoRGB / grayscaleFlowerFruitPanicle / ear / spikeLeafStem / branchWhole plant / canopy / plot / fieldClassificationObject detection

Abstract Observer bias and inconsistencies in traditional plant phenotyping methods limit the accuracy and reproducibility of fine-grained plant analysis. To address these limitations, TomatoMAP is introduced as a comprehensive dataset for Solanum lycopersicum . The dataset contains 68,080 RGB images: 3,616 high-resolution macrophotographs (3648 × 5472) with semantic annotations, and 64,464 moderate-resolution images (1080 × 1440) captured from 12 plant poses at four camera elevations. Each image is accompanied by manually annotated bounding boxes for seven regions of interest (leaves, panicle, flower clusters, fruit clusters, axillary shoot, shoot, and whole-plant area) and by labels spanning 50 BBCH classes representing phenologically growth stages. A general cascading structure is proposed. For real-time applicability, models emphasizing the accuracy-efficiency trade-off (MobileNetv3, YOLOv11, and Mask R-CNN) are prioritized and benchmarked against multiple state-of-the-art models. Performance is assessed using accuracy, mAP, inference FPS, and normalized confusion matrices. In a study involving five domain experts, AI models trained on TomatoMAP achieves comparable accuracy levels. Reliability of automated fine-grained phenotyping is supported by Cohen’s Kappa statistics and inter-rater agreement heatmaps.

Why it matches plant phenotyping methodsトマトの多視点画像、器官領域・生育ステージ注釈を備えたデータセットを構築し、画像モデルの精度・効率・専門家一致度をベンチマークしており、植物フェノタイピング手法が中心である。

titleTomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping.
Reproduction assets foundThe paper's authors publicly release their analysis code (dataset construction scripts for TomatoMAP-Cls/Det and model training/evaluation code) on GitHub. The TomatoMAP phenotype image dataset itself is deposited at e!DAL (10.5447/ipk/2025/14), but no matching URL is present in the allowed list, so only the code asset
Code · publicThe scripts for constructing TomatoMAP-Cls and TomatoMAP-Det, as well as the code used for model evaluation, are available at: https://github.com/0YJ/TomatoMAP.Open asset ↗https://github.com/0YJ/TomatoMAPhtml-lines:423-479
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published24 Feb 2026Data in briefCited by 1 · OpenAlex ↗

TLS-grapevine2024: A terrestrial laser scanner point cloud dataset of grapevines at different phenological stages.

GrapevineField / plotLiDAR / point cloudRGB / grayscaleMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldGrowth / time-series analysisBiomass / plant weightGrowth / development / phenology

Grapevines ( Vitis vinifera L.) undergo structural and physiological changes throughout the growing season, progressing through distinct phenological stages that require regular monitoring. This dataset consists of high-resolution point cloud data acquired with a stationary terrestrial laser scanner (TLS) to document grapevine development from early leaf development to dormancy. Georeferenced point clouds were generated from 15 TLS scans along two vineyard rows at nine phenological stages. The dataset also includes multispectral and RGB photogrammetric point clouds and orthorectified raster products from an unmanned aerial vehicle survey conducted before harvest. Ground-truth measurements leaf area index, grape production, and pruning wood biomass were collected for each monitored grapevine. As a result, the dataset provides multi-temporal TLS observations that support grapevine structural analysis and development, phenological monitoring, and can be used for the development of AI-based models for precision viticulture.

Why it matches plant phenotyping methodsブドウの生育・構造・フェノロジーを対象とするTLS点群および関連画像データセットであり、植物フェノタイピング用の再利用可能なデータ基盤として中心的です。

titleTLS-grapevine2024: A terrestrial laser scanner point cloud dataset of grapevines at different phenological stages.
Reproduction assets foundThe paper is a Data in Brief article describing the TLS-grapevine2024 dataset itself, publicly deposited on Zenodo with DOI 10.5281/zenodo.16751663. This is a paper-specific, openly available asset containing the TLS point clouds, UAV imagery/rasters, and ground-truth agronomic measurements (LAI, grape production, prun
Dataset · publicditions: clear sky. Data source location Institution: University of Trás-os-Montes e Alto Douro City/Town/Region: Arroios, Vila Real, Norte Country: Portugal Coordinates: 41°17′28.83″N 7°43′17.90″W, Altitude: 435 m Data accessibility Repository name: Zenodo Data identification number: 10.5281/zenodo.16751663 Direct URL to data: https://doi.org/10.5281/zenodo.16751663 Related research article None 1. Value of the Data • This dataset covers nine phenological stages of grapevine growth from April 2024 to January 2025, providing multi-temporal terrestrial laser scanner (TLS) observations for structural and phenological analysis. • It includes TLS point clouds collected at multiple stages and muOpen asset ↗Zenodo · 10.5281/zenodo.16751663lines:1-50
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published21 Feb 2026Data in briefCited by 0 · OpenAlex ↗

A field-acquired RGB-Depth image dataset for computer vision-based baby broccoli detection and size estimation under varying illumination conditions.

Brassica vegetablesField / plotLiDAR / point cloudRGB / grayscaleRGB-D / ToFWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detection2D/3D reconstructionSegmentation

This data article describes a curated RGB-Depth image dataset captured using an Intel RealSense D435 stereo depth camera mounted on an autonomous mobile platform during field deployments at commercial baby broccoli farms in Victoria, Australia. The dataset comprises 1759 paired RGB images (640 × 480 pixels) and corresponding 16-bit depth frames acquired under both daytime (natural sunlight) and night-time (LED illumination) conditions, designed to support research in agricultural computer vision and robotic harvesting. Images were selected from 39,765 raw acquisitions through a reproducible Python curation pipeline applying quality filtering (blur detection, brightness thresholds, corruption detection), perceptual hash-based duplicate removal, and manual review. The final dataset includes 924 daytime and 835 night-time image pairs containing baby broccoli plants at various growth stages. The dataset provides RGB camera intrinsic parameters and pixel-aligned depth maps to enable 3D point cloud reconstruction. Potential applications include developing deep learning models for crop detection and segmentation, validating depth-based size estimation methods, and benchmarking illumination-robust vision systems. All data and curation code are publicly available under a CC BY 4.0 license.

Why it matches plant phenotyping methodsRGB-Depth画像データセットの構築と再現可能なキュレーションを中心とし、作物検出に加えてサイズ推定という植物形質の評価・ベンチマークに利用できるため。

titleA field-acquired RGB-Depth image dataset for computer vision-based baby broccoli detection and size estimation under varying illumination conditions.
Reproduction assets foundThe paper is a data article describing a public Mendeley Data repository containing the authors' field-acquired RGB-D baby broccoli image dataset (1759 image pairs, ground truth diameter annotations, camera intrinsics, and curation/annotation code), directly reproducing the paper's phenotyping measurements and analysis
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/px5p6zdk6k.3 Direct URL to data: https://data.mendeley.com/datasets/px5p6zdk6k/3Open asset ↗Mendeley Data · 10.17632/px5p6zdk6k.3html-lines:95-155
Code / dataset availability confirmedCrossref · Europe PMC · checked 15 Sept 2026
Published20 Feb 2026Scientific DataCited by 0 · OpenAlex ↗

A comprehensive UK crop yield dataset incorporating satellite, weather, and soil type information

Field / plotWhole plant / canopy / plot / fieldYield / biomass estimationYield / yield components

Abstract Agricultural research increasingly relies on data-driven approaches for crop yield prediction that complement more established crop growth models, including machine learning techniques. However, these approaches rely on large training datasets. Here, we present the Crop Yields, Climate, Soils, and Satellites (CYCleSS) dataset, a large-scale crop yield dataset derived from precision yield data for 934 fields across England on which a variety of crops are grown. In addition, the data also contains satellite-derived remote sensing data, weather data, and data on soil type, all aligned at a grid resolution of 10 km. Weather data is available at a daily temporal resolution, satellite data at 5-day resolution, while crop yield data is available at yearly resolution. This effort has been made possible through careful anonymisation of the yield data while preserving the alignment with remote sensing, weather, and soil data. This data will be useful both to train machine learning models of yield prediction as well as to parameterize mechanistic crop growth models. Furthermore, the anonymisation procedure itself will be of interest to the research community, as it represents a solution to a common problem on the interface of agricultural research and farming practice.

Why it matches plant phenotyping methods圃場単位の作物収量という植物形質を、衛星・気象・土壌情報と整合した再利用可能な大規模データセットとして構築しており、収量予測モデルの訓練・評価用データ基盤が中心です。

abstractHere, we present the Crop Yields, Climate, Soils, and Satellites (CYCleSS) dataset, a large-scale crop yield dataset derived from precision yield data for 934 fields across England
Reproduction assets foundThe paper's authors provide public R code for merging/aligning climate, soil, and Sentinel-1 data and anonymising yield data in a GitHub repository. The CYCLeSS dataset itself is on figshare, but that URL is not in the allowed list, so only the code asset is reported.
Code · publicnts of this repository. Researchers who are further interested in the underlying data should contact the authors affiliated with UKCEH. Code availability R code used to merge and align available UK climate, soil, and Sentinel-1 synthetic aperture radar data to the same 1 km 2 grid is provided in the following GitHub repository: https://github.com/alan-turing-institute/CYCLeSS-dataset-code . Dummy data and code needed to replicate the final process of merging climate, soil, and satellite data with UKCEH precision yield data and anonymisation of field locations is contained within the ‘CLYCESS_anonymisation.zip’ folder shared as part of this repository. R version 4.2.3 was used for the creatioOpen asset ↗https://github.com/alan-turing-institute/CYCLeSS-dataset-codelines:200-271
Code / dataset availability confirmedOpenAlex · Europe PMC · bioRxiv · checked 15 Sept 2026
Published19 Feb 2026bioRxiv (Cold Spring Harbor Laboratory)Cited by 0 · OpenAlex ↗

Transformers Outperform ConvNets for Root Segmentation: A Systematic Comparison Across Nine Datasets

RootSegmentationRoot system architecture

Abstract Root segmentation is a fundamental yet challenging task in image-based plant phenotyping. We present the first systematic comparison of Transformer and Convolutional Neural Network (ConvNet) architectures for root segmentation, evaluating 21 architectures across nine diverse datasets and comparing pre-trained models to training from scratch. Transformer-based models significantly outperform ConvNets for segmentation accuracy and root-diameter agreement. Pre-training significantly improves mean Dice from 0.623 to 0.666 ( p = 3.3 × 10 −10 ). We also find that Transformers benefit more from pre-training than ConvNets, with Dice improvements of +0.072 versus +0.022 ( p = 3.7 × 10 −4 ), supporting the hypothesis that fine-tuned Transformers transfer more effectively across large domain gaps. Among evaluated models, MobileSAM achieved the highest Dice score while maintaining computational efficiency. Dataset choice explained far more performance variance (70.9%) than model architecture (6.7%), suggesting that data curation matters more than model selection.

Why it matches plant phenotyping methods根の画像セグメンテーション手法を21種類・9データセットで体系比較し、植物フェノタイピングにおける精度と根径推定を検証しているため、方法評価が中心である。

abstractRoot segmentation is a fundamental yet challenging task in image-based plant phenotyping.
Reproduction assets foundThe paper's authors explicitly state that their training/segmentation analysis code is publicly available on GitHub. The nine root image datasets evaluated are cited prior public datasets (DeepRootLab, Grassland, Chicory, PRMI), not paper-specific assets of this study, so the authors' own code repository is the only in
Code · publicr of parameters, as these affect hardware requirements, running costs, and environmental impact. To jointly compare efficiency and accuracy, we ranked models by the mean of their Dice, parameter count, and FLOPs ranks, providing a simple combined metric for practitioners balancing these trade-offs. Training code is available at https://github.com/sotlampr/seg.Configuration selection To prevent overfitting to the test set, model selection used a two-stage procedure based on validation performance: Replicate selection: For each combination of model, dataset, learning rate, and pre-training, the replicate with the highest validation Dice was retained, along with its paired test result. HyperparOpen asset ↗sotlampr/segpdf-raw-page:4 lines:1-95
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published18 Feb 2026Scientific DataCited by 0 · OpenAlex ↗

FIP 1.0 soybean data: Insights on soybean growth from eight years of high-throughput image field phenotyping

SoybeanField / plotRGB / grayscaleWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenologyStress response / toleranceYield / yield components

Abstract Soybean growth is determined by the interaction of genetic, environmental, and management factors. In the context of future climate and climate extremes, understanding genotype by environment interaction (GxE) will be crucial for selecting resilient breeding lines and optimizing management practices to minimize stress. This requires an in depth elucidation of stressful weather conditions and differing temporal responses of genotypes to those conditions. In field studies, however, the environment is often treated as a static factor, and the specific effects of weather variability on crop growth remain poorly understood. Here, we present a longitudinal dataset comprising 17,247 high-resolution RGB images of soybean breeding lines collected throughout eight years in Eschikon, Switzerland. Top-of-canopy images were acquired throughout the entire growing seasons and complemented by hourly weather data, enabling a comprehensive analysis of soybean growth dynamics under varying field conditions. High spatio-temporal image resolution allows detailed analysis of growth dynamics and GxE, supporting identification of stress-tolerant genotypes to improve yield prediction and yield stability.

Why it matches plant phenotyping methods8年間の高スループット画像フェノタイピングによる大規模データセットを提示し、画像取得基盤と作物成長動態の解析を中心に扱っているため、方法論文として適格です。

titleFIP 1.0 soybean data: Insights on soybean growth from eight years of high-throughput image field phenotyping
Reproduction assets foundThe paper's canopy cover analysis code is publicly available on the authors' ETH GitLab repository. The FIP 1.0 soybean image/trait dataset itself is deposited in the ETH Research Collection and Hugging Face, but those URLs are not among the allowed URLs, so only the code asset qualifies.
Code · publicCode availability The code is available on: https://gitlab.ethz.ch/crop_phenotyping/fip-soybean-canopycover. Users with similar data can use the implemented workflow to get canopy cover from their experiments.Open asset ↗gitlab.ethz.ch/crop_phenotyping/fip-soybean-canopycoverhtml-lines:207-226
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published17 Feb 2026Plant phenomics (Washington, D.C.)Cited by 2 · OpenAlex ↗

Leaf-DETR: Progressive adaptive network with lower matching cost for dense leaves detection.

Field / plotLeafObject detection

Leaves are central indicators of photosynthesis and plant growth status, and their precise monitoring is crucial for smart agriculture. Dense leaf detection, as a foundation for leaf morphology analysis, must address challenges such as occlusion and overlap, directly enabling key tasks including phenotypic trait extraction, disease identification, and yield estimation. Leaves are the most important plant organs, and monitoring leaves is a crucial aspect of crop surveillance. Dense leaf detection plays an important role as a fundamental technology for leaf monitoring. Existing dense leaf detection methods rely on traditional modular detectors and generic feature extraction, lacking designs tailored to real-world dense leaf scenarios. The methods for dense leaf detection generally use traditional modular detectors and general feature extraction techniques, without designing methods specifically for dense leaves in reality. In detail, in complex field scenarios, it still faces challenges like incomplete individual feature extraction due to high leaf overlap and difficult network convergence caused by excessive leaf density. To this end, we propose the Leaf-DETR framework, which effectively addresses these challenges through the Progressive Feature Fusion Pyramid Network (P-FPN) and the Crowded Query Refinement Strategy (CQR). First, we construct the largest dense leaf detection dataset to date, containing 1696 images and 85,375 annotation boxes. Second, P-FPN alleviates the feature confusion problem of overlapping leaves through the multi-stage fusion of features and the Adaptive Feature Aggregation module (AFA), enhancing the interaction between low-level details and high-level semantics. Third, the CQR strategy significantly reduces the matching cost of crowded candidate boxes and improves the network convergence efficiency by culling a crowded query method and introducing a one-to-many matching mechanism. Finally, experimental results show that Leaf-DETR improves mAP@50 by 1% and AR@300 by 1.4% over the baseline model on our self-constructed dataset, outperforming existing detection methods. Furthermore, the model exhibits extremely fast training convergence and demonstrates strong generalization capability on both field-collected monitoring images and other staple crops, fully highlighting its practical value in complex agricultural scenarios. Finally, experiments show that Leaf-DETR outperforms existing detection methods on the self-built dataset and demonstrates good performance generalization in monitoring collected images, as well as for other staple food crops, which verifies its practicality in complex agricultural scenarios. The code and detailed information are available at http://leafdetr.samlab.cn.

Why it matches plant phenotyping methods葉の密集検出モデルとデータセットを開発・評価し、葉形態などの表現型抽出を可能にする画像ベース手法が研究の中心であるため。

abstractDense leaf detection, as a foundation for leaf morphology analysis, must address challenges such as occlusion and overlap, directly enabling key tasks including phenotypic trait extraction, disease identification, and yield estimation.
Reproduction assets foundThe paper's data availability statement explicitly points to an authors' public site (http://leafdetr.samlab.cn) hosting the Leaf-DETR code and detailed information, qualifying as a paper-specific public code asset. The self-constructed KiwiFruitLeaf dataset (1696 images, 85,375 annotation boxes) is described but its公开
Code · publicThe code and detailed information are available at http://leafdetr.samlab.cn . For testing purposes, detailed instructions for running the model can be found in the repository's README file.Open asset ↗lines:504-529
Code / dataset availability confirmedCrossref · Europe PMC · checked 15 Sept 2026
Published13 Feb 2026PlantsCited by 3 · OpenAlex ↗

Pepper-4D: Spatiotemporal 3D Pepper Crop Dataset for Phenotyping

Pepper / chilliField / plotLiDAR / point cloudWhole plant / canopy / plot / fieldClassificationObject detectionSegmentationTrackingGrowth / development / phenology

Pepper (Capsicum annuum) is a globally significant horticultural crop cultivated for its culinary, medicinal, and economic value. Traditional approaches for boosting the agricultural production of pepper, notably, expanding farmland, have become increasingly unsustainable. Recent advancements in artificial intelligence and 3D computer vision have started to transform crop cultivation and phenotyping, which has shed new light on increasing production by advanced breeding. However, currently, the field still lacks 3D pepper data that contains enough detail for organ-level analysis. Therefore, we propose Pepper-4D, a new, high-precision 4D point cloud dataset that records both the spatial structure and temporal development of pepper plants across various continuous growth stages. Our dataset is divided into three subsets, including a total of 916 individual point clouds from 29 indoor-cultivated pepper plant samples. Our dataset provides manual annotations at both the plant-level and organ-level, supporting phenotyping tasks such as pepper growth status classification, organ semantic segmentation, organ instance segmentation, organ growth tracking, new organ detection, and even the generation of synthetic 3D pepper plants.

Why it matches plant phenotyping methods植物の器官レベル表現型解析を支援する4D点群データセットを構築し、成長状態分類・器官分割・追跡などを可能にする研究であり、フェノタイピング用データ基盤が中心です。

abstractPepper-4D, a new, high-precision 4D point cloud dataset that records both the spatial structure and temporal development of pepper plants across various continuous growth stages.
Reproduction assets foundThe authors publicly release the Pepper-4D spatiotemporal 3D pepper point cloud dataset (with plant- and organ-level annotations) and associated code via a GitHub repository stated in the Data Availability Statement. CloudCompare is a generic third-party tool, not a paper-specific asset.
Dataset · public.J.; writing—original draft preparation, F.A.; writing—review and editing, D.L.; visualization, F.A. and D.L.; supervision, D.L.; project administration, H.Y.; funding acquisition, D.L. and H.Y. All authors have read and agreed to the published version of the manuscript. Data Availability Statement Data and code can be found at https://github.com/foysalahmed10/Pepper-4D (accessed on 9 February 2026). Conflicts of Interest The authors declare no conflicts of interest. Funding Statement This work was supported in part by the Shanghai Sailing Program under Grant 24YF2701200, in part by the Fundamental Research Funds for the Central Universities under Grant 2232025D-50, and in part by Donghua UnOpen asset ↗foysalahmed10/Pepper-4Dlines:238-264
Code · publicang Q., Zeng Y., Hou J., Zhe X. WarpingGAN: Warping multiple uniform priors for adversarial 3D point cloud generation; Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; New Orleans, LA, USA. 18–24 June 2022; pp. 6397–6405. Associated Data Data Availability Statement Data and code can be found at https://github.com/foysalahmed10/Pepper-4D (accessed on 9 February 2026).Open asset ↗foysalahmed10/Pepper-4Dlines:308-314
Code / dataset availability confirmedCrossref · checked 13 Sept 2026
Published11 Feb 2026Scientific DataCited by 2 · OpenAlex ↗

Terrestrial and Airborne Laser Scanning Dataset of Trees in the Shivalik Range, India with Field Measurements and Leaf–Wood Classifications

Field / plotLiDAR / point cloudRGB / grayscaleLeafStem / branchWhole plant / canopy / plot / fieldClassificationSegmentation

Abstract Annotated datasets are essential for training and evaluating machine learning models in forest ecology. This dataset provides high-resolution, annotated LiDAR point clouds of 674 individual trees from 12 forest plots in the Shivalik Range of northern Haryana, India, representing 24 species. Data were acquired using Terrestrial Laser Scanning (TLS) and Airborne Laser Scanning (ALS), include field-measured attributes such as species identity and Diameter at Breast Height (DBH), and terrestrial and aerial RGB imagery. TLS point clouds were georeferenced and co-registered with centimetre-level accuracy, enabling precise integration with ALS data. The dataset includes segmented individual trees and wood–leaf classifications, suitable for applications such as tree morphology analysis, biomass estimation, and species classification. To support benchmarking, outputs from established classification algorithms (LeWoS, TLSeparation, CANUPO, and Random Forest) are included. As one of the first open-access LiDAR datasets from Indian tropical forests, it provides critical reference data for developing and validating forest structure models. It can also aid biomass mapping efforts in support of large-scale missions such as NASA-ISRO’s NISAR and ESA’s BIOMASS.

Why it matches plant phenotyping methods個体樹木のLiDAR点群・RGB画像と樹木セグメンテーションを含む公開データセットで、樹形解析や森林構造モデルの開発・検証、分類アルゴリズムのベンチマークを目的としており、植物形質取得が中心です。

abstractThis dataset provides high-resolution, annotated LiDAR point clouds of 674 individual trees from 12 forest plots in the Shivalik Range of northern Haryana, India, representing 24 species.
Reproduction assets foundThe paper's authors explicitly state that all code used for data processing, wood-leaf classification, feature extraction, and tree volume estimation is openly available on GitHub at https://github.com/moonis-ali/Dataset, which is an allowed URL. The paper's core LiDAR dataset is deposited on Zenodo (10.5281/zenodo.153
Code · publicAll code used for data processing, wood-leaf classification, feature extraction, and tree volume estimation is openly available on GitHub at https://github.com/moonis-ali/Dataset .Open asset ↗https://github.com/moonis-ali/Datasetlines:479-553
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published9 Feb 2026Scientific dataCited by 13 · OpenAlex ↗

A Large-Scale In-the-wild Dataset for Plant Disease Segmentation.

Field / plotSegmentationStress / disease detectionDisease symptoms / severity

Plant diseases pose significant threats to agriculture, making proper diagnosis and effective treatment crucial for protecting crop yields. In automatic diagnosis processing, image segmentation helps to identify and localize diseases. Developing robust image segmentation models for detecting plant diseases requires high-quality annotations. Unfortunately, existing datasets rarely include segmentation labels and are typically confined to controlled laboratory settings, which fail to capture the complexity of images taken in the wild. Motivated by these, we established a large-scale segmentation dataset for plant diseases, dubbed PlantSeg. In particular, PlantSeg is distinct from existing datasets in three key aspects: (1) Annotation types: PlantSeg includes detailed and high-quality disease area masks. (2) Image sources: PlantSeg primarily comprises in-the-wild plant disease images rather than laboratory images provided in existing datasets. (3) Scale: PlantSeg contains the largest number of in-the-wild plant disease images, including 7,774 diseased images with corresponding segmentation masks. This dataset provides an ideal yet unified benchmarking platform for developing advanced plant disease segmentation algorithms.

Why it matches plant phenotyping methods植物病害領域の画像セグメンテーション用データセットを構築し、病害領域マスクとベンチマーク基盤を提供することが中心で、植物の病害状態を直接推定する方法論的貢献である。

abstractThis dataset provides an ideal yet unified benchmarking platform for developing advanced plant disease segmentation algorithms.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicThe codes for the baseline reproduction are presented in https://github.com/tqwei05/PlantSeg.Open asset ↗https://github.com/tqwei05/PlantSeghtml-lines:720-764
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published7 Feb 2026DataCited by 0 · OpenAlex ↗

In Situ Crop and Soil Data and UAV Imagery from Winter Wheat Fields in a Bulgarian Site

WheatAerial / UAVField / plotWhole plant / canopy / plot / fieldBiomass / plant weightDisease symptoms / severityLeaf traitsPhotosynthesis / fluorescencePigment / colour / senescencePlant / canopy height

This data descriptor presents a dataset comprising crop and soil parameters measured in winter wheat fields near the town of Knezha, Bulgaria. The data were collected as part of a project evaluating the potential of vegetation indices derived from Sentinel-2 satellite imagery to predict biophysical and biochemical crop parameters. The core dataset consists of measurements obtained from 20 m × 20 m field plots and includes a broad range of parameters: leaf area index, fraction of absorbed photosynthetically active radiation, vegetation cover fraction, chlorophyll content, above-ground biomass, plant nitrogen content, biological yield, surface soil moisture, spectral reflectance, plant density, crop height, visual assessments of disease or pest damage, and data on weed occurrence. The dataset is complemented by unmanned aerial vehicle imagery, crop calendars, and field management information. The main soil types in the study area were characterized through soil profiles, while meteorological data were obtained from an automated weather station. The data were collected during the 2016–2017 and 2017–2018 agricultural seasons. The dataset is freely available for download and serves as a valuable resource for researchers in remote sensing—particularly for validating satellite-derived products—as well as for specialists involved in winter wheat monitoring, modeling, and agronomic studies.

Why it matches plant phenotyping methods冬小麦の複数の植物形質を含む再利用可能なデータセットを提示し、UAV画像や衛星由来指標の検証を主目的としているため、植物フェノタイピング用データセットとして採用。

abstractThis data descriptor presents a dataset comprising crop and soil parameters measured in winter wheat fields near the town of Knezha, Bulgaria.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicDataset: In situ and UAV dataset with crop and soil parameters obtained from winter wheat fields. https://doi.org/10.5281/zenodo.17475742.Open asset ↗zenodo · 10.5281/zenodo.17475742pdf-page:1 lines:1-56
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 15 Sept 2026
Published6 Feb 2026Plant PhenomicsCited by 1 · OpenAlex ↗

Fine-grained 3D rice phenotyping via multi-scale NeRF and multimodal segmentation.

RiceField / plotMultimodalNeRF / 3D Gaussian SplattingLiDAR / point cloudSeed / grainWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionSegmentation

Fine-grained 3D phenotypic analysis of rice plays a vital role in rice breeding and yield estimation. However, a comprehensive rice data acquisition and segmentation pipeline is still lacking. While Neural Radiance Fields (NeRF) have shown impressive results in crop-level 3D reconstruction, their high sensitivity to data volume and camera viewpoints often leads to reconstruction failures for rice. In addition, the large-scale rice point clouds, coupled with heavy occlusion and visual similarity among grains, pose significant challenges for fine-grained trait extraction. To address the challenge of reconstructing rice point clouds under low-quality data conditions, we propose a novel method named Multi-Scale NeRF(MSNeRF). This method incorporates a structure-detail collaborative reconstruction mechanism and a dynamic initialization density scheduling strategy. Furthermore, we introduce a multimodal and multitask rice dataset (MMR) as a benchmark resource for future research. For rice point cloud segmentation, we develop Vision Rice Knowledge Graph Network(VRKGNet), which comprises an image segmentation module, a projection module, and a point cloud segmentation module enhanced with a Transformer to enlarge the receptive field. VRKGNet performs standalone point cloud segmentation and integrates image segmentation results from multiple viewpoints as prior knowledge to enhance semantic and instance-level segmentation. Extensive experiments demonstrate that MSNeRF achieves high-fidelity point cloud reconstruction with as few as 10 viewpoints. VRKGNet achieves superior rice plant segmentation with a semantic segmentation mIoU of 88.79% and an instance segmentation AP 25 of 84.55%, outperforming mainstream algorithms.

Why it matches plant phenotyping methods米の3D形質取得・再構成・分割を中核とする手法開発であり、データセット/ベンチマークも提供しているため、植物フェノタイピング手法文献に該当する。

abstractwe propose a novel method named Multi-Scale NeRF(MSNeRF)
Reproduction assets foundThe paper's authors explicitly state that the source code for MSNeRF and VRKGNet is publicly available on GitHub with testing scripts and test cases to reproduce the main results. The MMR dataset itself is only available upon request from the corresponding author, so it does not qualify as a public asset.
Code · publicof Hefei Artificial Intelligence Breeding Accelerator Co. Ltd. ( NB2024005-02 ). Data availability The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper: • MSNeRF : https://github.com/qfwysw/MSNeRF.git • VRKGNet : https://github.com/qfwysw/VRKGNet.git The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author. Declaration of competing interest The authors declare that they have no known competing financial iOpen asset ↗https://github.com/qfwysw/MSNeRF.gitlines:620-663
Code · publicator Co. Ltd. ( NB2024005-02 ). Data availability The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper: • MSNeRF : https://github.com/qfwysw/MSNeRF.git • VRKGNet : https://github.com/qfwysw/VRKGNet.git The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could haveOpen asset ↗https://github.com/qfwysw/VRKGNet.gitlines:620-663
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published5 Feb 2026Plant phenomics (Washington, D.C.)Cited by 0 · OpenAlex ↗

Dual-guided asymmetric MP-former for rice root instance segmentation.

RiceRootMorphology / geometry measurementSegmentationRoot system architecture

Root phenotypic traits such as length and number are critical indicators of plant growth and productivity. However, accurate extraction of these traits remains challenging due to the slender morphology, dense overlap, and frequent occlusion within root systems. Traditional digital image processing methods suffer from low throughput and limited robustness, while most deep learning-based approaches rely on semantic segmentation, which fails to distinguish individual roots and therefore limits their applicability in instance-level phenotypic analysis.To address these limitations, we propose Dual-Guided Asymmetric MP-Former (DGA-MP-Former), a novel instance segmentation model tailored for root phenotyping, with rice roots as a representative case. Building upon the MP-Former framework, our model introduces two key components: the Guided-Enhancement Pixel Decoder (GEPD) and the Asymmetric Dual-Query Decoder (ADQD). The GEPD enhances multi-scale feature representations via Hybrid Convolution Aggregator, Semantic-Guided Fusion Module and Frequency-Guided Feature Enhancement Module, effectively capturing fine root structures and low-contrast regions. ADQD employs asymmetric interaction between semantic and instance queries to improve long-range dependency modeling and instance separation in occluded scenarios.Additionally, we present the Rice Root Segmentation Dataset (RRSD), comprising of 343 high-resolution images with instance-level annotations. Experimental results show that DGA-MP-Former achieves state-of-the-art performance on RRSD, with 57.2% AP 0.5:0.95 and 87.4% AP 0.5 . Importantly, the accurate instance segmentation results enable reliable computation of instance-level geometric traits, such as root perimeter and area. To quantitatively assess phenotypic measurement accuracy, Relative Area Error (RAE) and Relative Perimeter Error (RPE) are further introduced, achieving 26.4% and 20.2%, respectively. These results demonstrate that the proposed method effectively bridges instance segmentation accuracy and phenotypic quantification reliability, supporting high-throughput and precise root phenotyping.

Why it matches plant phenotyping methodsイネ根の個体別セグメンテーションモデルを開発し、データセット提供、性能評価、および根の形態形質推定まで行っており、植物フェノタイピング手法が研究の中心である。

abstractwe propose Dual-Guided Asymmetric MP-Former (DGA-MP-Former), a novel instance segmentation model tailored for root phenotyping
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicThe Rice Root Segmentation Dataset is open sourced for the research community at ”https://github.com/Run-19/DGA-mpformer”.Open asset ↗Run-19/DGA-mpformerhtml-lines:442-469
Code / dataset availability confirmedOpenAlex · Crossref · checked 14 Sept 2026
Published4 Feb 2026Remote SensingCited by 1 · OpenAlex ↗

Unsupervised Tree Detection from UAV Imagery and 3D Point Clouds via Distance Transform-Based Circle Estimation and AIC Optimization

Aerial / UAVLiDAR / point cloudRGB / grayscaleWhole plant / canopy / plot / fieldObject detection

This work proposes a novel tree detection methodology, named DTCD (Distance Transform Circle Detection), based on a fast circle detection method via Distance Transform and Akaike Information Criterion (AIC) optimization. More specifically, a visible-band vegetation index (RGBVI) is calculated to enhance canopy regions, followed by morphological filtering to delineate individual tree crowns. The Euclidean Distance Transform is then applied, and the local maxima of the smoothed distance map are extracted as candidate tree locations. The final detections are iteratively refined using the AIC to optimize the number of trees with respect to canopy coverage efficiency. Additionally, this work introduces DTCD-PC, a modified algorithm tailored for point clouds, which significantly enhances detection accuracy in complex environments. This work makes a significant contribution to tree detection in the following ways: (1) by creating a tree detection framework entirely based on an unsupervised technique, which outperforms state-of-the-art unsupervised and supervised tree detection methods; (2) by introducing a new urban dataset, named AgiosNikolaos-3, that consists of orthomosaics and photogrammetrically reconstructed 3D point clouds, allowing the assessment of the proposed method in complex urban environments. The proposed DTCD approach was evaluated on the Acacia-6 dataset, consisting of UAV images of six-month-old Acacia trees in Southeast Asia, demonstrating superior detection performance compared to existing state-of-the-art techniques, both unsupervised and supervised. Additional experiments were conducted in the custom-developed Urban Dataset, confirming the robustness and generalizability of the DTCD-PC method in heterogeneous environments.

Why it matches plant phenotyping methodsUAV画像・3D点群から個体樹冠を抽出する新規手法を開発し、複数データセットで精度・頑健性を評価しているため、植物形態の取得・抽出が中心である。

abstractThis work proposes a novel tree detection methodology, named DTCD (Distance Transform Circle Detection), based on a fast circle detection method via Distance Transform and Akaike Information Criterion (AIC) optimization.
Reproduction assets foundThe authors state that the MATLAB code implementing the DTCD/DTCD-PC method, together with the datasets (Acacia-6, AgiosNikolaos-3) and results, is publicly available at their project page. Since the article is published (accepted), this is an actionable public asset containing the paper's tree-detection analysis code,
Code · public.P.; All authors have read and agreed to the published version of the manuscript. Funding: This research received no external funding. Data Availability Statement: The code implementing the proposed method together with our results, and the links to the datasets are publicly available after paper acceptance at the following linkhttps://sites.google.com/site/costaspanagiotakis/research/tree-detection-dtcd, accessed on 30 January 2026. Conflicts of Interest: The authors declare no conflicts of interest. Abbreviations The following abbreviations are used in this manuscript: AIC Akaike Information Criterion AMS3D Adaptive Mean Shift 3D CHM Canopy Height Model CHT Circular Hough Transform CSP ComOpen asset ↗pdf-layout-page:24 lines:1-62
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published3 Feb 2026Data in briefCited by 0 · OpenAlex ↗

An open image dataset of Indonesian soybean seed varieties (Anjasmoro, Grobogan, DEGA-1) for agricultural research and machine learning applications.

SoybeanLaboratory / benchtopSeed / grainSegmentationFruit / seed / panicle traits

Soybean ( Glycine max L. ) performs an important position as a main resource of protein in Indonesia. Its quality and productivity can be assessed based on the characteristics of its seed. Accordingly, the identification process through the observation of soybean seed traits is a crucial step in plant breeding and quality assurance. Manual approaches rely on manual observation, which is subjective, prone to human error and time-consuming. With the improvement of artificial intelligence, automated seed identification has appeared as a potential solution. However, progress is constrained by the lack of open and standardized image datasets, especially for locally bred varieties in developing countries. To address this gap, we propose an open image dataset of Indonesian soybean seeds from three widely cultivated and plant-bred varieties: Anjasmoro, Grobogan, and DEGA-1. The dataset consists of high-resolution seed images captured with an Epson L360 flatbed scanner, with the optical resolution fixed at 800 dots per inch, yielding images of 6800 × 9359 pixels. All raw images are saved in JPG format. No manually segmentation masks are released in this version, instead of using Deeplab V3+ with MobileNet as backbone to enable the automated seed image segmentation. The curated dataset is intended to support a broad range of applications, including computer vision tasks such as image classification and segmentation, as well as research in plant breeding, seed quality assessment, and agricultural informatics. By providing a standardized and publicly accessible resource, this dataset contributes to the advancement of interdisciplinary studies at the intersection of agriculture and artificial intelligence.

Why it matches plant phenotyping methods大豆種子画像を標準化して公開するデータセット研究であり、種子形質の自動画像解析・セグメンテーションを支援する方法論的資源が中心です。

titleAn open image dataset of Indonesian soybean seed varieties (Anjasmoro, Grobogan, DEGA-1) for agricultural research and machine learning applications.
Reproduction assets foundThe paper is a data descriptor for a public Mendeley Data repository containing the authors' own soybean seed image dataset (raw scans and segmented seed images) used for seed phenotyping, with an explicit direct URL and DOI.
Dataset · publicData accessibility Repository name: Mendeley Data Data identification number: DOI: 10.17632/c733bjz4m3.3 Direct URL to data: https://data.mendeley.com/datasets/c733bjz4m3/3Open asset ↗Mendeley Data · 10.17632/c733bjz4m3.3html-lines:115-142
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published29 Jan 2026Data in briefCited by 5 · OpenAlex ↗

Agri-vision Bangladesh: A multi-crop augmented image dataset for automated disease diagnosis in Bottle Gourd, Zucchini, Papaya, and Tomato.

TomatoField / plotRGB / grayscaleLeafStress / disease detectionDisease symptoms / severity

This article introduces Agri-Vision Bangladesh, a comprehensive, augmented image dataset designed to advance automated disease diagnosis in four economically vital agricultural crops: Bottle Gourd ( Lagenaria siceraria ), Zucchini ( Cucurbita pepo ), Papaya (Carica papaya), and Tomato ( Solanum lycopersicum ). Addressing the scarcity of region-specific agricultural data, a total of 5266 original images were acquired directly from diverse agricultural fields in Bangladesh using a SONY ALPHA 7 II full-frame camera under natural lighting conditions. The dataset encompasses 28 distinct classes, covering a wide spectrum of biotic stressors including viral (Mosaic Virus, Leaf Curl), fungal (Downy Mildew, Anthracnose, Alternaria Blight), bacterial (Bacterial Blight, Xanthomonas), and pest-induced damage (Insect Hole, White Spot), alongside Healthy samples. To ensure scientific reliability, each image underwent a rigorous two-stage validation process by senior agronomists. To tackle class imbalance and facilitate the training of data-intensive Deep Learning models, the dataset was expanded using a Python-based augmentation pipeline incorporating geometric transformations (rotation, flipping) and photometric adjustments (noise, brightness) resulting in a final repository of 28,000 images (5266 original and 22,734 augmented). All files are standardized to 512×512 pixels in JPG format. This expert-validated resource serves as a critical benchmark for developing robust computer vision algorithms (e.g., CNNs, Vision Transformers) for precision agriculture, enabling research into fine-grained classification, object detection, and cross-crop transfer learning in subtropical farming environments.

Why it matches plant phenotyping methods植物病害症状を画像で分類するための専門家検証済みデータセットを構築し、再利用可能なベンチマークとして提供しているため、植物表現型取得法が中心です。

abstractThis article introduces Agri-Vision Bangladesh, a comprehensive, augmented image dataset designed to advance automated disease diagnosis
Reproduction assets foundThe paper is a Data in Brief article describing the Agri-Vision Bangladesh multi-crop leaf disease image dataset, publicly deposited on Mendeley Data with an explicit direct URL and DOI (10.17632/8t6k37ztxc.2). This is a paper-specific public asset containing the original and augmented plant images used in the study.
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/8t6k37ztxc.2 Direct URL to data: https://data.mendeley.com/preview/8t6k37ztxc?a=a88a48f1-a9b0-4354-a081-cc8f1e936364Open asset ↗Mendeley Data · 10.17632/8t6k37ztxc.2html-lines:93-117
Code / dataset availability confirmedCrossref · Europe PMC · checked 14 Sept 2026
Published23 Jan 2026Scientific DataCited by 1 · OpenAlex ↗

High-Resolution Leaf Image Sequences with Geometric Alignment for Dynamic Phenotyping of Foliar Diseases.

WheatRGB / grayscaleLeafImage / point-cloud registrationSegmentationGrowth / time-series analysisDisease symptoms / severity

Abstract Time-resolved phenotyping of disease symptoms enables dissection of resistance mechanisms and improves diagnosis, but acquiring phenotypic data at satisfactory scale remains challenging. Advances in imaging and image processing have improved measurement precision, robustness, and throughput, but further improvements are needed for practical application. We present a data set comprising 12,520 high-resolution (~0.03 mm/pixel) RGB images representing 1,032 time series of wheat leaves with developing disease symptoms. All images are geometrically aligned with a median precision of 0.16 mm (≈5 pixels). The dataset includes transformation matrices, symptom segmentation masks, metadata on treatments, weather, crop phenology, and disease occurrence, and a lightweight Python toolkit for loading, aligning, inspecting, and editing image sequences. These resources enable detailed investigation of leaf-level disease dynamics such as lesion, pustule, and fruiting body emergence rates, lesion growth, and dynamic interactions of disease development with spatial and environmental contexts. They offer a broad basis for developing improved methods for image alignment and symptom detection, segmentation, and tracking, possibly by tackling these connected challenges within a single end-to-end framework.

Why it matches plant phenotyping methods葉の病徴を対象とした高解像度時系列画像データセットで、幾何位置合わせ、病徴セグメンテーション、追跡用ツールを提供しており、植物病害表現型の取得・解析基盤が中心である。

abstractWe present a data set comprising 12,520 high-resolution (~0.03 mm/pixel) RGB images representing 1,032 time series of wheat leaves with developing disease symptoms.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicWe provide a lightweight Python toolkit to facilitate loading, inspection, and curation of the image sequences and their associated processing products in the associated Git repository (https://github.com/and-jonas/sympathique-wheat).Open asset ↗github.com/and-jonas/sympathique-wheathtml-lines:317-337
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 5 Sept 2026
Published22 Jan 2026Frontiers in Plant ScienceCited by 0 · OpenAlex ↗

DCSFormer: a high-precision method for cotton seedling point cloud organ segmentation.

CottonLiDAR / point cloudLeafStem / branchWhole plant / canopy / plot / fieldSegmentation

Introduction: Accurately segmenting cotton seedling organs from 3D point clouds is fundamental for high-throughput plant phenotyping and digital breeding. However, cotton seedling segmentation remains challenging due to fine-scale and complex organ morphology, uneven point density with noise, and the lack of high-quality annotated datasets. Methods: To address these issues, we propose DCSFormer, a tailored extension of Point Transformer V3 designed for cotton seedling point cloud segmentation. The model introduces the DCS Block, which leverages dynamic sparse expert routing and dual-channel attention to adaptively capture global semantic dependencies and subtle local geometric variations, thereby improving stem-leaf boundary discrimination. In addition, the proposed CLFSkip replaces traditional skip connections with a cross-layer fusion strategy, effectively integrating multi-scale features while preserving organ-level details. We also constructed an annotated cotton seedling dataset to support training and evaluation. Results and Discussion: Experimental results show that DCSFormer achieves 93.67% mIoU, 95.83% mPrec, 97.35% mRec, and 96.56% mF1, outperforming multiple comparison models. Furthermore, when evaluated against baseline models on two public datasets, Crops3D and Pheno4D, DCSFormer exceeds the baseline across all four metrics, further validating its effectiveness and generalizability. This work provides an effective solution for precise cotton seedling organ segmentation.

Why it matches plant phenotyping methods綿花幼苗の3D点群から器官を抽出する手法を開発し、アノテーション済みデータセットの構築と複数データセットでの性能検証を行っており、植物表現型取得が中心である。

abstractAccurately segmenting cotton seedling organs from 3D point clouds is fundamental for high-throughput plant phenotyping and digital breeding.
Reproduction assets foundThe authors constructed an annotated cotton seedling point cloud dataset (100 samples with semantic/instance organ labels and ground-truth traits) used for training and evaluating DCSFormer, and the data availability statement points to a public Kaggle repository containing it. No author analysis code or trained model/
Dataset · publicThe datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: https://www.kaggle.com/datasets/tengfeiliu333/dcsformer-cotton/croissant/download .Open asset ↗Kaggle · tengfeiliu333/dcsformer-cottonlines:957-1015
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published21 Jan 2026Scientific reportsCited by 4 · OpenAlex ↗

Generalizability and transferability of machine learning models using hyperspectral reflectance data for maize traits.

MaizeMultispectral / hyperspectralLeafMorphology / geometry measurementPhysiological trait estimationPhotosynthesis / fluorescence

Hyperspectral reflectance provides rapid, non-destructive phenotyping of plant leaves. These data have been used to develop machine learning models for predicting diverse plant traits, yet key challenges remain. We collected hyperspectral reflectance data together with 25 anatomical, gas exchange, and chlorophyll fluorescence traits from 320 recombinant inbred lines grown over three seasons. Using these data, we systematically (1) compare the performance of PLSR and SVR across a wide range of traits, including also slow fluorescence kinetics, (2) assess model generalizability and transferability, and (3) investigate how different aggregation strategies affect predictive accuracy. Based on a nested cross-validation framework, single cross-validation with MSE as metric performed comparably to repeated cross-validation or PRESS-based calibration. Optimal performance of trait-specific predictions was found to be dependent on the combination of model and data aggregation levels. Structural and biochemical traits showed the best generalizability and transferability, whereas physiological traits, particularly those derived from gas exchange and fluorescence kinetics, exhibited markedly reduced transferability. Together, these results provide a rigorous benchmark for evaluating machine learning models for trait prediction from hyperspectral reflectance data, and highlight both the opportunities and limitations for achieving robust generalization across diverse environments and genotypes.

Why it matches plant phenotyping methodsハイパースペクトル反射データから植物形質を予測する機械学習手法を、複数形質・環境・遺伝子型で系統的に比較し、一般化性と転移性を厳密にベンチマークしているため、方法論が中心である。

abstractHyperspectral reflectance provides rapid, non-destructive phenotyping of plant leaves.
Reproduction assets foundThe paper's Data availability statement explicitly deposits all code and raw hyperspectral/trait data in a public GitHub repository, matching an allowed URL.
Code · publicAll code and raw data to ensure reproducibility of the results can be accessed at: [https://github.com/Rudan-X/HyperspectralML](https:/github.com/Rudan-X/HyperspectralML).Open asset ↗Rudan-X/HyperspectralMLlines:158-246
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published21 Jan 2026Plant methodsCited by 0 · OpenAlex ↗

The Tonoplast Topology Index-a new metric for describing vacuole organization.

ArabidopsisLaboratory / benchtopMicroscopyCell / cellular structureRootMorphology / geometry measurement

Background The plant vacuole arises by orchestrated interplay of membrane trafficking, cytoskeletal rearrangements and a variety of signaling pathways. In the root, the characteristic large central vacuole develops by endomembrane reorganization occurring mainly in the transition zone. The vacuole's bounding membrane-the tonoplast-can be visualized in vivo using fluorescent protein markers, allowing for quantitative analysis of confocal microscopy images. Tonoplast organization can thus serve as a sensitive indicator of changes to any of the processes involved in vacuole biogenesis. The Vacuolar Morphology Index (VMI) is widely accepted as a quantitative measure of vacuole structure. However, this metric has two drawbacks-it only reflects the size of the largest vacuolar compartment (missing therefore possible differences in the organization of smaller compartments), and its determination is labor intensive, limiting its use on large datasets. Results We developed an alternative metric for describing vacuole organization, named the Tonoplast Topology Index (TTI), which overcomes the above-mentioned shortcomings of the VMI. We compared the performance of our protocol with VMI on a simulated dataset and on real data. To validate the methods´ performance, we used it to confirm the previously reported differences in vacuole shape and size between Arabidopsis thaliana roots grown on the surface of an agar medium compared to those embedded inside the agar. Both VMI and TTI could efficiently detect the relatively subtle changes in vacuole organization depending on the position of the root in the agar, and provided correlated results. However, only TTI produced data with close to normal value distribution, simplifying subsequent statistical evaluation. Conclusions We present the protocol for TTI determination as a two-stage semi-automated procedure involving microscopic image analysis employing an ImageJ macro and subsequent processing of numeric data in the Jupyter Notebook environment, together with benchmarking image data. Since this implementation is freeware-based, platform-independent and (relatively) user-friendly, we hope it will find its use as a high throughput, added value alternative to the VMI metric.

Why it matches plant phenotyping methods植物液胞構造を定量化する新規指標と半自動画像解析プロトコルを開発し、シミュレーションおよび実画像で既存指標と比較・検証しているため、植物フェノタイピング手法が中心である。

abstractWe developed an alternative metric for describing vacuole organization, named the Tonoplast Topology Index (TTI)
Reproduction assets foundThe paper deposits its benchmark confocal image dataset in the EMBL-EBI BioImage Archive (S-BIAD2226) and its TTI analysis software (ImageJ macro and Jupyter/Python scripts) on GitHub, both with explicit public availability statements.
Dataset · publicImage data generated and analyzed in the current study are available in the EMBL-EBI BioImage Archive repository, accession number S-BIAD2226Open asset ↗EMBL-EBI BioImage Archive · S-BIAD2226lines:141-163
Code · publicArchive copy, additional sample data and possible future updates of the software tool generated here are also available at [ https://github.com/GeorgeCaldarescu/TTI-Tonoplast-Topology-Index ] .Open asset ↗GitHub · GeorgeCaldarescu/TTI-Tonoplast-Topology-Indexlines:141-163
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published21 Jan 2026Data in briefCited by 0 · OpenAlex ↗

BrinjalFruitX: A field-collected image dataset for machine learning and deep learning-based disease identification in brinjal fruits.

Eggplant / aubergineField / plotFruitClassificationDisease symptoms / severity

Brinjal (Solanum melongena) or eggplant is one of the four most essential vegetable crops that are grown in Bangladesh and contribute significantly to the agricultural industry of the country. Brinjal supports the livelihood of numerous small farmers; however, brinjal is severely susceptible to various fruit diseases, which have serious impacts on yield quality and may cause considerable economic losses. While most existing plant disease datasets primarily focus on leaf-related disorders, only a limited number include fruit-related diseases and even those contain very few classes. This gap is significant because fruit diseases directly affect crop quality, market value, and overall yield. This is why we present here a new and comprehensive dataset that is unparalleled, exclusively for brinjal fruit diseases. This data set consists of 1823 high-quality, labelled images, across five distinct classes: Phomopsis Blight, Shoot and Fruit Borer, Fruit Cracking, Wet Rot, and Healthy Fruit. The images were collected from real farm conditions in numerous areas of Bangladesh to ensure a robust sample of varied environmental and farming practices impacting the growth of diseases. This dataset is designed with the unique aim to support plant disease research and enhance training of deep learning models for autonomous disease detection. Lastly, the dataset will allow early disease detection, enhancing crop management practice, reduction of losses, and increasing farmers' economic returns. The release of this dataset will encourage agricultural research as well as practical use in precision agriculture.

Why it matches plant phenotyping methodsナス果実の病徴を対象とした画像データセットの構築・公開が中心であり、植物の病害状態を画像から識別する再利用可能な表現型データ資源に該当する。

abstractwe present here a new and comprehensive dataset that is unparalleled, exclusively for brinjal fruit diseases.
Reproduction assets foundThe paper's brinjal fruit disease image dataset (1823 labeled images, five classes) is publicly deposited on Mendeley Data, and the authors' model training/augmentation code is publicly available on GitHub.
Code · publicThe complete code, along with augmentation scripts and model development, is publicly available in our GitHub repository [12].Open asset ↗GitHubhtml-lines:299-357
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published19 Jan 2026Journal of integrative plant biologyCited by 2 · OpenAlex ↗

Stem microanatomical phenomic uncovers a potential role for ZmLSM2 in regulating maize stem bending strength.

MaizeX-ray / CTStem / branchMorphology / geometry measurementArchitecture / morphology / geometryStress response / tolerance

Modern maize stems possess a well-developed vascular bundle system, which is critical for providing mechanical support and lodging resistance. However, characterization of the microanatomical features of vascular bundles and their functional implications in stem mechanics remains challenging, primarily due to technical limitations in high-throughput microanatomical analysis of stem tissues. We thus constructed data sets consisting of over 500,000 maize stem CT images from a maize diversity panel of 383 inbred lines. We evaluated 32 microanatomical phenotypes of maize basal internodes across two environments in different years. By incorporating engineering mechanics parameters, we calculated novel characteristics of the vascular bundles, including the moment of area (MOA) and the polar moment of inertia (PMOI). Through the high-density phenotypic data set, we identified multiple stem microanatomical phenotypes strongly associated with lodging resistance, particularly of vascular bundle mechanical traits. By integrating population genetic profiling, we discovered and confirmed that ZmLSM2 (U6 small nuclear ribonucleoprotein specific Sm-like 2) serves as a key regulator of stem mechanical strength, might function in RNA processing and maturation within vascular stem cells, identifying novel genetic targets for improving maize lodging resistance. This approach demonstrates the value of combining advanced phenotyping with multi-omics analyses for crop improvement. These discoveries will deepen the understanding of plant stem biomechanical principles and provide novel targets for enhancing lodging resistance in crop breeding programs.

Why it matches plant phenotyping methodsトウモロコシ茎のCT画像から微細構造形質を高スループットに抽出する表現型解析基盤とデータセットが研究の中心であり、単なる生物学的測定ではない。

abstracttechnical limitations in high-throughput microanatomical analysis of stem tissues
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicCT cross‐section images of the third internode from 383 maize inbred lines grown in Beijing and Sanya during two growing seasons can be downloaded via the link: https://pan.baidu.com/s/1CP2kkAmTvy1zi3QJGtKSWQ?pwd=JIPB . Extraction code: JIPB.Open asset ↗lines:204-306
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published16 Jan 2026PloS oneCited by 2 · OpenAlex ↗

Towards practical AI for agriculture: A self-supervised attention framework for Spinach leaf disease detection.

SpinachLeafClassificationStress / disease detectionDisease symptoms / severity

Malabar spinach is a nutrient-dense leafy vegetable widely cultivated and consumed in Bangladesh. Its productivity is often compromised by Alternaria leaf spot and straw mite infestations. This work proposes an efficient and interpretable deep learning framework for automatic Malabar spinach leaf disease classification. A curated dataset of Malabar spinach images collected from Habiganj Agricultural University and supplemented with public samples was categorized into three classes: Alternaria, straw mite, and healthy leaves. A lightweight SpinachCNN established a strong baseline, while Spinach-ResSENet, enhanced with squeeze-and-excitation modules, improved channel-wise attention and feature discrimination. A customized Vision Transformer (SpinachViT) and SwinV2-Base were further investigated to assess the benefits of transformer-based architectures under limited data. To mitigate annotation scarcity, we employed SimSiam-based self-supervised pretraining on unlabeled images, followed by supervised fine-tuning with cross-entropy or a hybrid objective combining cross-entropy and supervised contrastive loss. The best-performing domain-optimized model, SimSiam-CBAM-ResNet-50, incorporated Convolutional Block Attention Modules and achieved 97.31% test accuracy, 0.9983 macro ROC-AUC, and low calibration error, while maintaining robustness to Gaussian and salt-and-pepper noise. Although a SwinV2-Base benchmark pretrained on ImageNet-22k reached slightly higher accuracy (97.98%, 98.99% with test-time augmentation), its 86.9M parameters and reliance on large-scale pretraining reduce feasibility for edge deployment. In contrast, the SimSiam-CBAM model offers a more parameter-efficient and deployment-friendly solution for real-world agricultural applications. Model decisions are interpretable via Grad-CAM, Grad-CAM++, and LayerCAM, which consistently highlight biologically relevant lesion regions. The spinach dataset used in this study is publicly available on: https://huggingface.co/datasets/saifullah03/malabar_spinach_leaf_disease_dataset.

Why it matches plant phenotyping methods葉画像から病害状態を分類する深層学習手法を開発・比較し、公開データセットと解釈可能性・頑健性も評価しており、植物表現型取得が中心である。

abstractThis work proposes an efficient and interpretable deep learning framework for automatic Malabar spinach leaf disease classification.
Reproduction assets foundThe paper's Malabar spinach leaf disease image dataset (the phenotyping input used for all measurements) is explicitly stated as publicly available on Hugging Face, with the URL given in the abstract and Data Availability Statement. No author analysis code or trained model checkpoints are explicitly deposited.
Dataset · publicThe spinach dataset used in this study is publicly available on: https://huggingface.co/datasets/saifullah03/malabar_spinach_leaf_disease_dataset.Open asset ↗huggingface · saifullah03/malabar_spinach_leaf_disease_datasethtml-lines:585-614
Code / dataset availability confirmedCrossref · checked 5 Sept 2026
Published13 Jan 2026Earth System Science DataCited by 3 · OpenAlex ↗

Global near real-time 500 m 10 d FPAR dataset from MODIS and VIIRS for operational agricultural monitoring and crop yield forecasting

Whole plant / canopy / plot / fieldCalibration / preprocessingGrowth / time-series analysisPhotosynthesis / fluorescenceYield / yield components

Abstract. Climate change and extreme weather events pose challenges to food security, emphasizing the need for reliable and timely monitoring of crop and rangeland conditions. For this purpose, long-term consistent Earth Observation datasets on vegetation conditions are typically used in early warning and crop yield forecast systems. However, the near-real-time (NRT) production of high quality datasets and the need to guarantee long-term records present various challenges. To address these, we present a NRT global dataset of Fraction of Photosynthetically Active Radiation (FPAR) at 500 m resolution, optimized for agricultural applications. Our dataset combines MODIS-FPAR (Collection 6.1) and VIIRS-FPAR (Collection 2) data, ensuring continuity from 2000 to well beyond 2030. We applied a robust filtering approach based on the Whittaker smoother to produce reliable FPAR estimates in NRT, accounting for sparse and irregular spaced observations due to cloud cover. The dataset is composed of two 10 d filtered timeseries: (1) MODIS-FPAR for 2000 to 2023, being the reference dataset, and (2) intercalibrated VIIRS-FPAR for 2018 onward. While several methods can effectively smooth and gap-fill FPAR data (i.e., using observations before and after the estimation date), our method is designed for optimal filtering in NRT (i.e., using only prior observations). Our approach yields six successive estimates of the same FPAR data point with increasing quality: an inital estimate immediately after the 10 d reference period, four subsequent estimates every 10 d using new observations, and a final consolidated estimate 90 d later. The implemented filtering ingests the available FPAR observations and their original quality assessment (QA) layers. To avoid unrealistic extrapolation when observations are sparse, we impose constraints, season and location specific, to FPAR estimates. We then intercalibrated the VIIRS-FPAR with the MODIS-FPAR filtered timeseries, using a mean difference correction approach, to ensure consistency between both series. This paper describes the filtering and intercalibration method used, the quality assessment of resulting timeseries, and details the obtained products and the corresponding QA layers. The NRT FPAR dataset is publicly available through the Joint Research Centre Data Catalogue, https://doi.org/10.2905/1aac79d8-0d68-4f1c-a40f-b6e362264e50 (Seguini et al., 2025).

Why it matches plant phenotyping methodsMODIS/VIIRSから植物キャノピー状態であるFPARを推定するNRTフィルタリング・相互較正手法とデータセットの開発、品質評価が中心であり、単なる農業モニタリングへの routine measurement ではない。

abstractThis paper describes the filtering and intercalibration method used, the quality assessment of resulting timeseries, and details the obtained products and the corresponding QA layers.
Reproduction assets foundThe paper describes its own global NRT 500 m 10 d filtered FPAR dataset (MODIS and intercalibrated VIIRS timeseries with QA layers), explicitly stated to be publicly and freely available via the JRC Data Catalogue DOI and directly downloadable from the ASAP server, with visualization in the ASAP Warning Explorer. This衍
Dataset · publicThe NRT FPAR dataset is publicly available through the Joint Research Centre Data Catalogue, https://doi.org/10.2905/1aac79d8-0d68-4f1c-a40f-b6e362264e50 ( Seguini et al. , 2025 ) .Open asset ↗10.2905/1aac79d8-0d68-4f1c-a40f-b6e362264e50lines:158-173
Dataset · publicor can be directly downloaded from the following server https://agricultural-production-hotspots.ec.europa.eu/data/MO6_FPAR/ (last access: 30 September 2025).Open asset ↗lines:245-257
Code / dataset availability confirmedOpenAlex · checked 5 Sept 2026
Published9 Jan 2026Earth system science dataCited by 1 · OpenAlex ↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity

Field / plotMultispectral / hyperspectralLeafVisualization / data managementLeaf traitsPhotosynthesis / fluorescence

Abstract. Accurate assessment of leaf functional traits is crucial for a diverse range of applications from crop phenotyping to parameterizing global climate models. Leaf reflectance spectroscopy offers a promising avenue to advance ecological and agricultural research by complementing traditional, time-consuming gas exchange measurements. However, the development of robust hyperspectral models for predicting leaf photosynthetic capacity and associated traits from reflectance data has been hindered by limited data availability across species and environments. Here we introduce the Global Spectra-Trait Initiative (GSTI), a collaborative repository of paired leaf hyperspectral and gas exchange measurements from diverse ecosystems. The GSTI repository currently encompasses over 7500 observations from 397 species and 41 sites gathered from 36 published and unpublished studies, thereby offering a key resource for developing and validating hyperspectral models of leaf photosynthetic capacity. The GSTI database is developed on GitHub (https://github.com/plantphys/gsti, last access: 4 January 2026) and published to ESS-DIVE https://doi.org/10.15485/2530733, Lamour et al., 2025). It includes gas exchange data, derived photosynthetic parameters, and key leaf traits often associated with traditional gas exchange measurements such as leaf mass per area and leaf elemental composition. By providing a standardized repository for data sharing and analysis, we present a critical step towards creating hyperspectral models for predicting photosynthetic traits and associated leaf traits for terrestrial plants.

Why it matches plant phenotyping methods葉のハイパースペクトルとガス交換・光合成形質を標準化して収録するデータベースを構築し、植物フェノタイピングモデルの開発・検証に供することが中心である。

abstractHere we introduce the Global Spectra-Trait Initiative (GSTI), a collaborative repository of paired leaf hyperspectral and gas exchange measurements from diverse ecosystems.
Reproduction assets foundThe paper describes the GSTI database of paired leaf hyperspectral and gas-exchange measurements, with both the data and R processing/model-fitting code publicly available on GitHub and archived releases on ESS-DIVE.
Code · publicThe GSTI data and code are available in the public GitHub repository at https://github.com/plantphys/gsti (last access: 4 January 2026)Open asset ↗https://github.com/plantphys/gstilines:537-549
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published8 Jan 2026Scientific DataCited by 4 · OpenAlex ↗

The Multi-Sensor and Multi-Temporal Dataset of Multiple Crops for In-Field Phenotyping and Monitoring

Aerial / UAVField / plotLiDAR / point cloudRGB / grayscaleMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldImage / point-cloud registrationBiomass / plant weightLeaf traits

Abstract Phenotyping is crucial for understanding crop trait variation and advancing research, but is currently limited by expensive, labor-intensive monitoring. New phenotypic trait monitoring methods are being proposed to reduce this so-called phenotyping bottleneck via automation. These methods are often data-driven, requiring a dataset recorded with a specific sensor and corresponding reference values for developing novel methods. To this end, we present the MuST-C (Multi-Sensor, multi-Temporal, multiple Crops) dataset, which contains field data from various sensors collected over a growing season, covering six crop species. All data was georeferenced for alignment across sensors and dates. To collect our dataset, we deployed aerial and ground robotic platforms equipped with RGB cameras, LiDARs, and multispectral cameras, aiming to capture a wide variety of modalities and observations from different viewpoints. In addition to sensor data, we also provide manually collected leaf area index and biomass reference measurements. Our dataset enables the development of novel automatic phenotypic trait estimation methods, allows comparisons across different sensors, and generalizability across crop species.

Why it matches plant phenotyping methods複数センサー・ロボットプラットフォームによる圃場フェノタイピング用データセットを構築・提供し、形質推定法の開発、センサー比較、汎化評価を可能にすることが中心的な貢献である。

abstractwe present the MuST-C (Multi-Sensor, multi-Temporal, multiple Crops) dataset
Reproduction assets foundThe paper's MuST-C multi-sensor, multi-temporal crop phenotyping dataset (RGB/multispectral images, LiDAR point clouds, LAI and biomass reference measurements) is publicly available via the authors' project webpage, and the authors' custom Python processing/loading code is publicly available on GitHub.
Dataset · publicThe MuST-C dataset is available via our project webpage https://www.ipb.uni-bonn.de/data/MuST-C/or directly via the bonndata public access repository 10.60507/FK2/OX9XTM34Open asset ↗html-lines:421-440
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published8 Jan 2026Data in briefCited by 0 · OpenAlex ↗

Corn seed dataset based on hyperspectral and RGB images.

MaizeLaboratory / benchtopRGB / grayscaleMultispectral / hyperspectralSeed / grainClassificationCalibration / preprocessing

This study employed an HY-6010-S hyperspectral imaging system, covering a spectral range of 400-1000 nm, combined with an RGB industrial camera to acquire multimodal data. The dataset simulates phenotypic analysis scenarios of maize seeds under controlled laboratory conditions, with the ambient temperature maintained at 20-25°C. Comprehensive testing was conducted using 12 different maize varieties. Approximately 200 seed samples were collected per variety, resulting in a total sample size of about 2400, each subjected to hyperspectral and RGB image acquisition. Preprocessing steps included noise reduction, background removal, band selection, and modality alignment. To ensure the accuracy and reliability of the experimental data, HHIT software and Python were utilized for data processing. This dataset plays a significant role in seed variety classification, phenotypic analysis, precision agriculture, and machine learning applications.

Why it matches plant phenotyping methodsトウモロコシ種子のマルチモーダル画像を収集・前処理した再利用可能なデータセットであり、種子の表現型解析を主要目的としているため、フェノタイピング手法・データセット研究に該当する。

abstractThis study employed an HY-6010-S hyperspectral imaging system, covering a spectral range of 400-1000 nm, combined with an RGB industrial camera to acquire multimodal data.
Reproduction assets foundThe paper is a Data in Brief article depositing its own multimodal maize seed hyperspectral and RGB image dataset (2400 seeds, 12 varieties) on Mendeley Data, with a direct public URL and DOI given in the article.
Dataset · publicRepository name: Mendeley Data Data identification number: doi: 10.17632/4n4xbnx8sr.1 Direct URL to data: https://data.mendeley.com/datasets/4n4xbnx8sr/1Open asset ↗Mendeley Data · 10.17632/4n4xbnx8sr.1html-lines:1-110
Code / dataset availability confirmedarXiv · OpenAlex · checked 15 Sept 2026
Published1 Jan 2026arXivCited by 0 · OpenAlex ↗

CropNeRF: A Neural Radiance Field-Based Framework for Crop Counting

AppleCottonPearField / plotNeRF / 3D Gaussian SplattingFruitCounting2D/3D reconstructionSegmentation

Rigorous crop counting is crucial for effective agricultural management and informed intervention strategies. However, in outdoor field environments, partial occlusions combined with inherent ambiguity in distinguishing clustered crops from individual viewpoints poses an immense challenge for image-based segmentation methods. To address these problems, we introduce a novel crop counting framework designed for exact enumeration via 3D instance segmentation. Our approach utilizes 2D images captured from multiple viewpoints and associates independent instance masks for neural radiance field (NeRF) view synthesis. We introduce crop visibility and mask consistency scores, which are incorporated alongside 3D information from a NeRF model. This results in an effective segmentation of crop instances in 3D and highly-accurate crop counts. Furthermore, our method eliminates the dependence on crop-specific parameter tuning. We validate our framework on three agricultural datasets consisting of cotton bolls, apples, and pears, and demonstrate consistent counting performance despite major variations in crop color, shape, and size. A comparative analysis against the state of the art highlights superior performance on crop counting tasks. Lastly, we contribute a cotton plant dataset to advance further research on this topic.

Why it matches plant phenotyping methodsNeRFと3Dインスタンスセグメンテーションを用いて作物個体・器官数を推定する画像ベース表現型計測手法を開発・検証しており、方法が研究の中心である。

abstractwe introduce a novel crop counting framework designed for exact enumeration via 3D instance segmentation.
Reproduction assets foundThe paper contributes a public infield cotton plant dataset (8 plants, ~150 iPhone images each, ground-truth boll counts, SAM instance masks) and states that source code, dataset, and multimedia are available at the authors' public project page, which is an allowed URL. The spectacularai GitHub URL is a generic third-p
Dataset · publicthat incorporates crop visibility and mask consistency, enabling robustness against occlusions and annotation discrepancies. • We release a public infield cotton plant dataset designed for 3D rendering and cotton boll counting tasks. The source code, dataset, and multimedia material associated with this project can be found at https://robotic-vision-lab.github.io/cropnerf . II Related Work II-A Image-Based Techniques Image-based methods typically employ object detection to identify crops within images. For example, Chen et al. [ 4 ] utilized multiple convolutional neural networks (CNNs) to map input images to total fruit counts. Similarly, Häni et al. [ 5 ] formulated crop counting as a multOpen asset ↗lines:108-187
Code / dataset availability confirmedEurope PMC · checked 13 Sept 2026
Published24 Dec 2025Data in briefCited by 1 · OpenAlex ↗

3-dimensional surface geometry, optical properties dataset of Scots pine and Norway spruce shoots.

Field / plotPhotogrammetry / SfM / MVSLiDAR / point cloudMultispectral / hyperspectralLeaf2D/3D reconstructionArchitecture / morphology / geometry

Conifer shoots possess highly complex geometrical structures at a very fine spatial resolution. Accurately characterizing the full architecture of a conifer shoot, which influences how radiation is scattered, has proven challenging. Previous radiative transfer models for coniferous stands have represented these structures in a relatively simplified or coarse manner. This paper presents a dataset that can be used for up-scaling of needle to shoot optical properties and studying the influence of detailed three-dimensional (3D) structure of shoot to light scattering within tree crown. The dataset includes 3D structural information as well optical properties of needles and twigs for 27 shoots of two conifer species present in both locations (3 shoots per species and position in the crown) - Scots pine ( Pinus sylvestris L.) and Norway spruce ( Picea abies L. Karst. ). The samples were collected on 22nd April 2024 in Rájec, the Czech Republic and 17th September 2024 in Järvselja, Estonia. Subsequently blue light 3D photogrammetry scanning technique was used to obtain their high-resolution 3D point cloud representations. Reflectance and transmittance measurements of needles were obtained using a spectroradiometer and an integrating sphere. For each of these samples, the dataset comprises a photo of the sampled shoot, obtained 3D surface reconstruction, and optical properties of conifer needles and twigs (hemispherical-conical reflectance and transmittance factors) in the spectral range of 400-2000 nm. A detailed 3D representation of needle shoots, when combined with radiative transfer modeling, may offer a means to study and compensate for inaccuracies in the measurement of needle optical properties and to enhance the assessment of shoot scattering characteristics.

Why it matches plant phenotyping methods針葉樹シュートの3D構造をフォトグラメトリで取得し、光学特性とともに再利用可能なデータセットとして提供しているため、植物形態・構造の計測手法が中心です。

abstractThis paper presents a dataset that can be used for up-scaling of needle to shoot optical properties and studying the influence of detailed three-dimensional (3D) structure of shoot to light scattering within tree crown.
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository containing the paper's own phenotyping measurements: 3D surface geometry models (.obj) of Scots pine and Norway spruce shoots, sample photos (.jpg), and needle/twig optical property spectra (HCRF/HCTF, .csv, 400-2000 nm). The repository,
Dataset · publicRepository name: Mendeley Data identification number: 10.17632/h39f9t7fjg.1 Direct URL to data: https://data.mendeley.com/datasets/h39f9t7fjg/2Open asset ↗Mendeley · 10.17632/h39f9t7fjg.1lines:47-74
Code / dataset availability confirmedCrossref · checked 5 Sept 2026
Published23 Dec 2025PlantsCited by 0 · OpenAlex ↗

Real-Time Callus Instance Segmentation in Plant Tissue Culture Using Successive Generations of YOLO Architectures

LentilLaboratory / benchtopLeafTissueSegmentation

Callus induction is a complex procedure in plant organ, cell, and tissue culture that underpins processes such as metabolite production, regeneration, and genetic transformation. It is important to monitor callus formation alongside subjective evaluations, which require labor-intensive care. In this research, the first curated lentil (Lens culinaris) callus dataset for instance segmentation was experimentally generated using three genotypes as one data set: Firat-87, Cagil, and Tigris. Leaf explants were cultured on MS medium fortified with different concentrations of gross regulators of BA and NAA to induce callus formation. Three biologically relevant stages, the leaf stage, the green callus, and the necrosis callus, were produced. During this process, 122 high-resolution images were obtained, resulting in 1185 total annotations across them. The dataset was evaluated across four successive generations (v5/7/8/11) of YOLO deep learning models under identical conditions using mAP, Dice coefficient, Precision, Recall, and IoU, together with efficiency metrics including parameter counts, FLOPs, and inference speed. The results show that anchor-based variants (YOLOv5/7) relied on predefined priors and showed limited boundary precision, whereas anchor-free designs (YOLOv8/11) used decoupled heads and direct center/boundary regression that provided clear advantages for callus structures. YOLOv8 reached the highest instance segmentation precision with mAP50@0.855, while it matched the accuracy with greater efficiency and achieved real-time inference with 166 FPS.

Why it matches plant phenotyping methods植物組織培養におけるカルスの形成段階・壊死状態を画像からインスタンスセグメンテーションする手法、データセット、モデル比較を中心に扱っており、植物状態の取得・定量化が本研究の主要な技術貢献である。

titleReal-Time Callus Instance Segmentation in Plant Tissue Culture Using Successive Generations of YOLO Architectures
Reproduction assets foundThe paper's lentil callus image dataset with annotations (122 images, 1185 annotations) is publicly available on Roboflow Universe per the Data Availability Statement. The YOLOv5 GitHub link and Ultralytics docs are generic third-party libraries, not authors' analysis code, and the FAO link is a cited reference, so all
Dataset · publicThe dataset used in this study, including annotated images for callus detection, is publicly available and can be accessed at Roboflow Universe: https://universe.roboflow.com/yunus-7v2b5/callus-hug7d , accessed on 13 September 2025. This repository contains all images and annotations generated and analyzed during the current study.Open asset ↗Roboflow Universe · callus-hug7dlines:314-345
Code / dataset availability confirmedEurope PMC · bioRxiv · checked 15 Sept 2026
Published22 Dec 2025bioRxivCited by 1 · OpenAlex ↗

Petal to the metal: The slow road to automating large-scale phenology labeling for herbarium specimens

FlowerAnnotation / quality controlObject detectionGrowth / development / phenology

ABSTRACT Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found in-house expert-curated annotations were essential for producing reliable results. Expert validation found relatively strong accuracy for detecting present floral structures, but still had moderately high false negative rates. Applying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present. However, only 2.9 million of these contained complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts. Still, this dataset represents a large compilation of historical herbarium-derived phenology records available as a resource for the phenology community. We end by demonstrating how integrating these machine-labeled records into Phenobase, a publicly-available phenology database, expands taxonomic and temporal coverage for large-scale phenological analyses, and discuss remaining challenges and next steps.

Why it matches plant phenotyping methods植物標本画像から花の存在を自動検出し、精度検証と大規模な phenology データセット化を行う機械学習手法が研究の中心であるため、植物フェノタイピング手法として適格です。

abstractwe present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community.
Reproduction assets foundThe paper's Data Availability Statement explicitly deposits the ensemble models, training/validation/test images, training data and final ensemble output on Zenodo, and the analysis code on GitHub; machine-labeled records are also served via the public Phenobase portal. All are paper-specific, public, and actionable.
Dataset · publicors contributed to drafts and gave final 454 approval for publication. 455 456 Data Availability Statement 457 The ensemble data models and a corresponding JSON file with model metadata data are 458 housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the SupportinOpen asset ↗Zenodo · 17675089pdf-raw-page:19 lines:1-55
Dataset · publictps://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supporting Information section at 466 the end of the article. 467 Appendix S1. List of difficult-to-annotate genera and families removed from training and 468 downstream data. 469 Appendix S2. Table S1. Validation results for held-oOpen asset ↗Zenodo · 10.5281/zenodo.17675089pdf-raw-page:19 lines:1-55
Code · publiclity Statement 457 The ensemble data models and a corresponding JSON file with model metadata data are 458 housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supporting Information section at 466 the end of the article. 467 Appendix S1. List of difficult-to-annotaOpen asset ↗GitHub · rafelafrance/phenobasepdf-raw-page:19 lines:1-55
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 Dec 2025Scientific reportsCited by 3 · OpenAlex ↗

TomatoRipen-MMT: transformer-based RGB and NIR spectral fusion for tomato maturity grading.

TomatoGreenhouseMultimodalRGB / grayscaleMultispectral / hyperspectralFruitClassificationSegmentationGrowth / development / phenology

Computer vision and multispectral imaging have increasingly become essential tools in modern precision agriculture. Accurate ripeness assessment is critical for yield optimization, reducing post-harvest losses, and enabling automated harvesting systems. However, traditional RGB-based approaches struggle to differentiate subtle maturity changes, and existing solutions often fail under varying lighting, occlusion, or cultivar-specific conditions. To address these challenges, this study focuses on the integration of complementary spectral cues for reliable tomato ripeness evaluation. The work utilizes a curated RGB-NIR tomato dataset comprising 224 hyperspectral samples, processed into aligned multimodal image pairs with balanced ripeness categories.The proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues. The novelty of the methodology lies in the dynamic cross-attention mechanism, which learns inter-modal dependencies between RGB and NIR signals for enhanced ripeness interpretation. Performance metrics including accuracy, precision, recall, F1-score, mIoU, and AUC were used to comprehensively evaluate the system. Experimental results demonstrate that TomatoRipen-MMT significantly outperforms all baseline RGB-only, NIR-only, and fusion methods, achieving 94.8% classification accuracy and 82.6% mIoU. These findings establish the effectiveness of multimodal Transformers for robust, high-precision fruit maturity assessment in controlled and greenhouse environments.

Why it matches plant phenotyping methodsトマト果実の成熟度という植物器官の状態を、RGB・NIR画像融合とTransformerで推定する手法を開発・評価しており、フェノタイピング手法が中心です。

abstractThe proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues.
Reproduction assets foundThe paper's phenotyping analysis is built on a publicly available USDA/NAL hyperspectral tomato dataset, explicitly linked in the Data Availability statement with an exact URL match. No author code or model checkpoints are disclosed.
Dataset · publicThe dataset analyzed in this study is publicly available at the https://agdatacommons.nal.usda.gov/articles/dataset/Data_from_b_Hyperspectral_Imaging_Analysis_for_Early_Detection_of_Tomato_Bacterial_Leaf_Spot_Disease_b_/26046328.Open asset ↗26046328html-lines:1038-1053
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published17 Dec 2025PloS oneCited by 3 · OpenAlex ↗

Empirically calibrated simulations reveal the limits of phenotypic clustering algorithms for biodiversity assessment in data-scarce crops.

MilletWhole plant / canopy / plot / field

Clustering algorithms are widely used for phenotypic characterization and germplasm management, particularly in data-scarce crops such as neglected and underutilized species (NUS) that lack genomic resources. However, their performance under biologically realistic conditions remains poorly understood. Standard clustering methods commonly applied in crop research often assume distinct, isotropic, and homogeneous clusters, assumptions rarely satisfied in real-world phenotypic datasets. We developed a flexible and empirically calibrated simulation framework, using phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios. Our simulations integrated heterogeneous trait distributions (normal, gamma), strong inter-trait correlations (up to r = -0.84), heteroscedasticity, and moderate population structure (mean Pst = 0.16 ± 0.001, achieved through iterative calibration). Each scenario was replicated 100 times, with clustering accuracy evaluated using external (ARI, NMI) and internal (Silhouette, Davies-Bouldin) validation metrics under standardized conditions. The results revealed consistently poor algorithm performance under realistic conditions (e.g., ARI < 0.07), including for widely used methods in Neglected and Underutilized Species (NUS) research such as K-means, GMM, and PAM. Notably, conventional validation metrics failed to detect biologically meaningful structure revealed by geometric diagnostics, highlighting a critical methodological limitation. Performance markedly improved under idealized conditions, validating our simulation framework. These findings highlight the risk of overinterpreting clustering outputs from weakly structured phenotypic datasets and expose key limitations in current biodiversity analysis practices, particularly those guiding plant genetic resource conservation programs. We provide an open-source R-based diagnostic tool, with parameter specifications to assist practitioners in selecting reproducible and interpretable clustering approaches for germplasm management and biodiversity assessment in data-scarce crops.

Why it matches plant phenotyping methods植物の表現型データを対象に、クラスタリング手法を現実的な条件でベンチマークするシミュレーション枠組みとR診断ツールを開発しており、表現型解析手法が研究の中心である。

abstractWe developed a flexible and empirically calibrated simulation framework, using phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios.
Reproduction assets foundThe paper's Data Availability statement explicitly deposits the complete R simulation/clustering/evaluation script on Zenodo (DOI 10.5281/zenodo.15877863), a paper-specific, publicly actionable code asset. The empirical fonio trait data belong to a prior cited study (Bio et al.), not this paper, and supporting files/DO
Code · publicthe complete R script used to simulate phenotypic datasets, apply clustering algorithms, and compute evaluation metrics is publicly available on Zenodo: https://doi.org/10.5281/zenodo.15877863Open asset ↗Zenodo · 10.5281/zenodo.15877863lines:107-122
Code / dataset availability confirmedbioRxiv · checked 13 Sept 2026
Published16 Dec 2025bioRxiv

A 0.6-meter resolution canopy height and structure model for the contiguous United States

Aerial / UAVWhole plant / canopy / plot / fieldMorphology / geometry measurementArchitecture / morphology / geometryPlant / canopy height

Above-ground vertical structure is a critical variable for ecosystem monitoring, carbon accounting, and land management. However, the high cost and limited coverage of airborne lidar hinder its widespread application. To address this, we developed NAIP-CHM, a 0.6-meter resolution canopy height and structure model (CHM) covering the contiguous United States, derived from National Agriculture Imagery Program (NAIP) aerial imagery. Unlike forestry-specific models that exclude human-made features, NAIP-CHM characterizes the full vertical structure of the landscape including vegetation, buildings, and infrastructure. We utilized a U-Net convolutional neural network with attention mechanisms and environmental conditioning, training and validating the model with a peer-reviewed, publicly available dataset of 22.8 million co-registered NAIP imagery and lidar-derived CHM pairs, with stratified sampling to ensure robustness in open-canopy ecosystems. The model achieved a pixel-wise root mean square error (RMSE) of 2.28 meters and an r2 of 0.87. Forested sites alone produced an r2 of 0.82 and RMSE of 3.82 meters. We provide the dataset, source code, and cloud-based tools to enable broad application without requiring specialized computational resources.

Why it matches plant phenotyping methods植生を含む景観の樹冠高・構造を航空画像から推定するモデルを開発し、公開データセットで検証している。植物キャノピーの明示的な構造形質推定が中心だが、建造物等も含むため植物以外の構造も対象とする点には留意が必要。

abstractwe developed NAIP-CHM, a 0.6-meter resolution canopy height and structure model (CHM) covering the contiguous United States, derived from National Agriculture Imagery Program (NAIP) aerial imagery.
Reproduction assets foundThe paper's NAIP-CHM canopy height model, its CONUS 0.6 m dataset, trained weights, and full training/inference code are all publicly released with explicit availability statements and author-hosted URLs (Rangeland Analysis Platform server, GitHub, Zenodo, Colab notebook, Earth Engine app).
Dataset · publicFor bulk download, COGs and associated index files are available via HTTP from the Rangeland Analysis Platform server ( http://rangeland.ntsg.umt.edu/data/naip-chm/ ).Open asset ↗Rangeland Analysis Platform serverlines:76-83
Model / weights · publicThe source code, trained model weights, validation data, and auxiliary datasets required to reproduce the results are permanently archived in a Zenodo repository 31 .Open asset ↗Zenodolines:89-134
Code / dataset availability confirmedOpenAlex · arXiv · checked 6 Sept 2026
Published15 Dec 2025arXiv (Cornell University)Cited by 0 · OpenAlex ↗

LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping

Rapeseed / canolaRGB / grayscaleLeafTracking

High resolution phenotyping at the level of individual leaves offers fine-grained insights into plant development and stress responses. However, the full potential of accurate leaf tracking over time remains largely unexplored due to the absence of robust tracking methods-particularly for structurally complex crops such as canola. Existing plant-specific tracking methods are typically limited to small-scale species or rely on constrained imaging conditions. In contrast, generic multi-object tracking (MOT) methods are not designed for dynamic biological scenes. Progress in the development of accurate leaf tracking models has also been hindered by a lack of large-scale datasets captured under realistic conditions. In this work, we introduce CanolaTrack, a new benchmark dataset comprising 5,704 RGB images with 31,840 annotated leaf instances spanning the early growth stages of 184 canola plants. To enable accurate leaf tracking over time, we introduce LeafTrackNet, an efficient framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network. During inference, leaf identities are maintained over time through an embedding-based memory association strategy. LeafTrackNet outperforms both plant-specific trackers and state-of-the-art MOT baselines, achieving a 9% HOTA improvement on CanolaTrack. With our work we provide a new standard for leaf-level tracking under realistic conditions and we provide CanolaTrack - the largest dataset for leaf tracking in agriculture crops, which will contribute to future research in plant phenotyping. Our code and dataset are publicly available at https://github.com/shl-shawn/LeafTrackNet.

Why it matches plant phenotyping methods葉レベルの時系列追跡という植物表現型取得手法を開発し、専用ベンチマークデータセットで評価しているため、方法が中心である。

abstractTo enable accurate leaf tracking over time, we introduce LeafTrackNet, an efficient framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network.
Reproduction assets foundThe authors explicitly state that the CanolaTrack dataset (5,704 annotated RGB images of 184 canola plants), the LeafTrackNet code, and trained model weights are publicly available at their GitHub repository.
Code · publicOur code and dataset are publicly available at https://github.com/shl-shawn/LeafTrackNet.Open asset ↗shl-shawn/LeafTrackNet · LeafTrackNetpdf-page:1 lines:1-53
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published15 Dec 2025Plant phenomics (Washington, D.C.)Cited by 3 · OpenAlex ↗

MaizeField3D: A curated 3D point cloud and procedural model dataset of field-grown maize from a diversity panel.

MaizeField / plotLiDAR / point cloudLeafStem / branchWhole plant / canopy / plot / field2D/3D reconstructionSegmentationArchitecture / morphology / geometry

The development of artificial intelligence (AI) and machine learning (ML) based tools for 3D phenotyping, especially for maize, has been limited due to the lack of large and diverse 3D datasets. 2D image datasets fail to capture essential structural details such as leaf architecture, plant volume, and spatial arrangements that 3D data provide. To address this limitation, we present MaizeField3D (website), a curated dataset of 3D point clouds of field-grown maize plants from a diverse genetic panel, designed to be AI-ready for advancing agricultural research. Our dataset includes 1045 high-quality point clouds of field-grown maize collected using a terrestrial laser scanner (TLS). Point clouds of 520 plants from this dataset were segmented and annotated using a graph-based segmentation method to isolate individual leaves and stalks, ensuring consistent labeling across all samples. This labeled data was then used for fitting procedural models that provide a structured parametric representation of the maize plants. The leaves of the maize plants in the procedural models are represented using Non-Uniform Rational B-Spline (NURBS) surfaces that were generated using a two-step optimization process combining gradient-free and gradient-based methods. We conducted rigorous manual quality control on all datasets, correcting errors in segmentation, ensuring accurate leaf ordering, and validating metadata annotations. The dataset also includes metadata detailing plant morphology and quality, alongside multi-resolution subsampled point cloud data (100k, 50k, 10k points), which can be readily used for different downstream computational tasks. MaizeField3D will serve as a comprehensive foundational dataset for AI-driven phenotyping, plant structural analysis, and 3D applications in agricultural research.

Why it matches plant phenotyping methods3D点群の収集・分割・注釈・手続き型モデル化を中核とする、植物表現型解析向けの再利用可能なデータセットである。

abstractwe present MaizeField3D (website), a curated dataset of 3D point clouds of field-grown maize plants from a diverse genetic panel, designed to be AI-ready for advancing agricultural research.
Reproduction assets foundThe paper's own MaizeField3D dataset (1045 TLS point clouds, 520 segmented/annotated plants, metadata, STL/DAT procedural model outputs) is publicly available on Hugging Face, with a project website and public GitHub code for the procedural NURBS surface generation used in the analysis.
Dataset · publicThe MaizeField3D dataset is publicly available on the Hugging Face Datasets platform at https://huggingface.co/datasets/BGLab/MaizeField3D. It includes high-resolution point clouds, segmented plant models, metadata, and reconstructed outputs in STL and DAT formats.Open asset ↗BGLab/MaizeField3Dhtml-lines:343-354
Code · publicThe code for procedural NURBS surface generation used in this work is available at https://github.com/baskargroup/ProceduralMaize3D.Open asset ↗baskargroup/ProceduralMaize3Dhtml-lines:343-354
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published11 Dec 2025Data in briefCited by 6 · OpenAlex ↗

A comprehensive combined dataset on Hibiscus and Tea plant leaf disease images for classifications.

TeaLeafClassificationDisease symptoms / severity

In this study, we present a combined image dataset created from two distinct plant species: Hibiscus and Tea leaf. The dataset consists of high-resolution images of leaves from both species, captured using a SONY α7 II DSLR camera and a OnePlus 7T lubricant Tea Leaf dataset includes images categorized into five disease classes: Algal Leaf Spot, Brown Blight, Grey Blight, Red Leaf Spot, and Healthy, while the Hibiscus Leaf dataset includes images labeled across eight conditions, including citrus spot, fungal infection, mild edge damage, and healthy foliage. To ensure balanced representation and address class imbalances, extensive data augmentation techniques-such as flipping, rotation, zooming, shifting, noise addition, and brightness adjustment-were applied, resulting in a total of 1,413 combined original images and 13,000 augmented images. The ConvNextTiny deep learning model was fine-tuned on this combined dataset to classify the various leaf conditions, achieving an overall accuracy of 96%. This demonstrates the model's robust performance and high discriminatory power across the diverse set of leaf diseases and conditions. This experiment highlights the utility of combining multiple plant species into a single dataset and utilizing a lightweight yet effective model like ConvNextTiny for plant disease classification. The resulting dataset, along with the model and training scripts, is publicly available to facilitate further research in plant pathology, computer vision, and smart farming applications, enabling more accurate and efficient early-stage disease detection for both Hibiscus and Tea plants.

Why it matches plant phenotyping methods植物葉の病害・健全状態を画像から分類するデータセットを構築し、分類モデルで性能評価しているため、植物フェノタイピング手法・ベンチマークが中心です。

abstractwe present a combined image dataset created from two distinct plant species: Hibiscus and Tea leaf
Reproduction assets foundThe paper's combined Hibiscus and Tea leaf disease image dataset is publicly deposited on Mendeley Data (DOI 10.17632/5bzy89brkv.4), and the authors' augmentation/training scripts are on a public GitHub repository; both are paper-specific, public, and directly actionable.
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/5bzy89brkv.4 Direct URL to data: https://data.mendeley.com/datasets/5bzy89brkv/4Open asset ↗Mendeley Data · 10.17632/5bzy89brkv.4lines:1-46
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published4 Dec 2025Data in briefCited by 0 · OpenAlex ↗

Dataset accompanying "Investigating the limits of spectroscopy for the estimation of foliar N and P in apple": Hyperspectral reflectance, foliar nutrient concentrations and associated metadata.

AppleGrowth chamberMultispectral / hyperspectralLeafPhysiological trait estimation

This dataset was generated to support research investigating the use of hyperspectral reflectance for the estimation of foliar nitrogen (N) and phosphorus (P) concentrations in apple ( Malus domestica ) trees. This article and the dataset it describes accompany an original research article submitted to Computers and Electronics in Agriculture entitled "Investigating the limits of spectroscopy for the estimation of foliar N and P in apple" [1]. Data were collected from a controlled potted experiment involving 150 'Golden Delicious' apple trees grown under varying nutrient supply regimes, including full nutrient supply, nitrogen- and phosphorus-deficient treatments, and trees infected with ' Candidatus Phytoplasma mali'. The experiment was conducted over the 2023 growing season at the Laimburg Research Centre in South Tyrol, Italy. All data and the accompanying code for its analysis is freely available in the associated GitHub repository [2]. Spectral data were collected using the Spectral Evolution SR-3500 field spectroradiometer with an attached leaf clip, producing high-resolution hyperspectral reflectance profiles (350-2500 nm) from the adaxial surface of fully expanded leaves. A total of 1189 leaf spectra were recorded and were matched to chemically analysed leaf samples. Corresponding foliar N and P concentrations (and others) were determined through laboratory analysis using the Dumas combustion method for nitrogen and ICP-OES following acid digestion for phosphorus. The dataset includes metadata detailing tree treatments, sampling dates, infection status, and shoot growth metrics. Additionally, R scripts used for data processing, spectral pre-treatment (including multiplicative scatter correction and Savitzky-Golay derivatives), feature selection (VIP and mRMR), and model development are provided. The dataset is suitable for reuse in the development and benchmarking of spectral models for nutrient estimation, especially in the context of field-based or remote sensing applications in horticulture. Its wide range of foliar nutrient values, inclusion of multiple physiological stresses, and detailed documentation make it a valuable resource for researchers working in precision agriculture, plant phenotyping, chemometrics, and hyperspectral data analysis.

Why it matches plant phenotyping methodsリンゴ葉のN・P濃度という植物生理形質を対象に、ハイパースペクトル反射データ、化学分析値、前処理・モデル開発コードを含む再利用可能なデータセットであり、植物フェノタイピング手法の開発・ベンチマークに直接資する。

abstractThis dataset was generated to support research investigating the use of hyperspectral reflectance for the estimation of foliar nitrogen (N) and phosphorus (P) concentrations in apple
Reproduction assets foundThe authors publicly release the paper's own hyperspectral leaf spectra (.sed files), matched foliar N/P concentrations, metadata, and R analysis scripts via a GitHub repository (also archived with Zenodo DOI 10.5281/zenodo.15600557), with explicit public availability and no registration required.
Dataset · publicData accessibility Repository name: Github Data identification number: DOI 10.5281/zenodo.15600557 Direct URL to data: https://github.com/HyperspectralCameron/Investigating-the-Limits-of-Spectroscopy-for-the-Estimation-of-Foliar-N-and-P-in-Apple.gitInstructions for accessing these data: All data and code are publicly available through the GitHub repository listed above. The repository includes raw spectral files (.sed), metadata files, and R scripts for pre-processing, modelling, and visualisation. No registration or authentication is required.Open asset ↗GitHub · DOI 10.5281/zenodo.15600557html-lines:84-123
Code · publicAll data and the accompanying code for its analysis is freely available in the associated GitHub repository [2].Open asset ↗GitHubhtml-lines:1-83
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published2 Dec 2025Biodiversity data journalCited by 0 · OpenAlex ↗

Dataset on flammability and functional traits of woody plants in a pine-oak forest of western Mexico.

Field / plotLeafStem / branchMorphology / geometry measurementLeaf traitsStress response / toleranceWater status / transpiration

Background Plant functional traits provide key information about species' ecological strategies and their responses to environmental disturbances such as fire. This dataset documents 14 morpho-functional traits of leaves (specific leaf area, leaf water content and leaf dry matter content), stems (maximum height, bark thickness, diameter at 40 cm, wood density, stem water content and stem dry matter content), one regenerative trait (resprouting capacity), as well as fire-related traits (ignition time, flaming time and flammability) and growth form in 50 woody plant species (27 trees, 22 shrubs and one liana) inhabiting a pine-oak forest in the "Barranca del Cupatitzio" National Park (BCNP), located in Uruapan, Michoacán, Mexico. This dataset is formatted according to the Darwin Core Archive standard and is publicly available for use. New information This dataset is standardised under the Darwin Core framework. It includes 14 morpho-functional and fire-related traits. The data were obtained from 50 woody species with a diameter at breast height (DBH) > 2.5 cm (27 trees, 22 shrubs and one liana), in a pine-oak forest located in the western Trans-Mexican Volcanic Belt, in the Municipality of Uruapan, Michoacán, Mexico. Here, we report flammability-related traits for these species for the first time. The collection of biological material and the measurement of functional traits followed internationally recognised protocols, ensuring methodological consistency and facilitating integration with other global datasets. The dataset includes values for flammability, ignition time, flaming time, specific leaf area, wood density, stem water and dry matter content, bark thickness, leaf water and dry matter content, maximum height, stem diameter at 40 cm above the ground, plant growth form and resprouting capacity. This information is particularly valuable for studies in functional ecology, ecological restoration, the dynamics of woody plant communities and fire management in temperate, fire-prone ecosystems.

Why it matches plant phenotyping methods植物の形態・機能・火災関連形質を体系的に収集し、Darwin Coreで標準化した再利用可能なデータセットであり、形質測定とデータ提供が中心である。

abstractThis dataset documents 14 morpho-functional traits of leaves
Reproduction assets foundThe paper is a data paper whose own trait/flammability dataset is deposited publicly on GBIF via DOI 10.15468/46f8xe, explicitly linked as the data package for this study's measurements.
Dataset · publiche Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits use, distribution and reproduction in any medium, provided the original authors are properly credited. Data resources Data package title Functional traits related to fire in woody species from Barranca del Cupatitzio National Park Resource link https://doi.org/10.15468/46f8xe Number of data sets 2 Data set 1. Data set name occurrence.txt Data format Darwin Core Data set 1. Column label Column description id Unique identifier for each occurrence. institutionID The identifier for the institution having custody of the specimens. institutionCode Full name of the institution having custody of the specimeOpen asset ↗10.15468/46f8xelines:87-297
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published28 Nov 2025Data in briefCited by 1 · OpenAlex ↗

A comprehensive image dataset of jute diseases.

Field / plotLeafClassificationDisease symptoms / severity

This Data Descriptor presents the Jute Diseases Image Dataset; a curated collection of 1390 high-resolution images aimed at supporting the development of machine learning models for timely identification and accurate diagnosis of jute (Corchorus) plant diseases. The dataset is categorized into five classes: Dieback (300), Holed (300), Mosaic (240), Stem Soft Rot (270), and Fresh (280) representing healthy leaves. Images were captured under varied natural lighting and directional conditions across diverse jute cultivation areas to enhance model generalizability. A rigorous pre-processing pipeline was applied, including uniform resizing to 1024 × 1024 pixels and removal of duplicate images to ensure data integrity. The dataset is organized into two components: a raw, pre-processed set and an augmented train-test split version, enabling immediate use in machine learning workflows. Additionally, Grad-CAM and Guided Grad-CAM techniques were applied to sample images to visualize and validate model attention on disease-relevant regions. This resource addresses the lack of labelled jute disease imagery and supports timely disease management, particularly for stakeholders in Bangladesh and other major jute-producing regions.

Why it matches plant phenotyping methods植物病害症状を画像として収集・ラベル化したデータセットであり、病害状態の画像ベース表現型判定を支援することが中心です。

abstractThis Data Descriptor presents the Jute Diseases Image Dataset; a curated collection of 1390 high-resolution images aimed at supporting the development of machine learning models for timely identification and accurate diagnosis of jute (Corchorus) plant diseases.
Reproduction assets foundThe paper's own jute disease image dataset (1390 labeled images, raw and augmented train/test splits) is publicly deposited in Harvard Dataverse with an explicit DOI and direct URL, matching an allowed URL. No separate analysis code or trained model checkpoint is publicly released.
Dataset · publicData accessibility Repository name: Harvard Dataverse Data identification number: https://doi.org/10.7910/DVN/FJ1DM1 Direct URL to data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FJ1DM1Open asset ↗Harvard Dataverse · doi:10.7910/DVN/FJ1DM1html-lines:1-91
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published25 Nov 2025Scientific dataCited by 2 · OpenAlex ↗

A long-term dataset of maize phenology observations from agrometeorological stations in Northeast China (1981-2024).

MaizeField / plotWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenology

We present a meticulously curated, long-term (1981-2024) dataset documenting maize phenology dynamics across Northeast China, the nation's most critical commercial grain base. Derived from 61 national agrometeorological stations, it captures the timing of 10 pivotal phenological stages (sowing, emergence, three-leaf, seven-leaf, jointing, tasseling, flowering, silking, milking, maturity) and derives the durations of 4 growth period lengths (sowing-jointing, jointing-silking, silking-maturity, sowing-maturity). The dataset underwent a rigorous, multi-tiered quality control protocol, including automated checks for internal consistency and expert arbitration for ambiguous records, ensuring high integrity. Subsequent analysis employed kernel density estimation to characterize the probability distribution of phenological events and univariate linear regression to quantify decadal trends. The resulting repository is substantial, comprising 976 georeferenced diagnostic plots in JPEG format and two primary data tables in XLSX format, with a total volume of 601.04 MB. Systematically organized by province and station, this dataset serves as a foundational empirical resource for quantifying climate-driven shifts in crop development, enhancing the parameterization and validation of process-based crop models, and informing the development of optimized cultivation practices and regional climate adaptation frameworks.

Why it matches plant phenotyping methodsトウモロコシの複数の生育ステージと生育期間を長期・広域に収録し、品質管理済みデータセットとして構築しているため、植物フェノタイピングデータセットが研究の中心です。

abstractit captures the timing of 10 pivotal phenological stages
Reproduction assets foundThe paper is a data descriptor whose maize phenology dataset (1981–2024, 61 stations, 10 phenological stages, diagnostic plots and XLSX tables) is openly deposited in Science Data Bank. No custom code was created.
Dataset · publich stage timing and duration, it empowers farmers and 307 agricultural planners to optimize production systems in response to evolving climatic 308 conditions, thereby enhancing regional food security resilience. 309 Data Availability 310 The dataset generated during this study is openly available in the Science Data Bank at 311 https://doi.org/10.57760/sciencedb.28709 or https://cstr.cn/31253.11.sciencedb.28709.312 Code availability 313 No custom code was created for the production of this dataset. 314 References 315 1.Cai C, Ding T, Chen W. (2024). Potential yield of world maize under global warming based on ARIMA-TR model. 316 Journal of Agrometeorology, 2024, 26(1). 317 2.Li M. RetrospectOpen asset ↗Science Data Bank · 10.57760/sciencedb.28709pdf-raw-page:15 lines:1-94
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published22 Nov 2025Data in briefCited by 0 · OpenAlex ↗

Phenology and health of Stenocereus Queretaroensis : A multimodal dataset combining multispectral imagery and spectrophotometry.

Field / plotMultimodalMultispectral / hyperspectralRaman / spectroscopyWhole plant / canopy / plot / fieldCalibration / preprocessingGrowth / development / phenology

This data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo), a native species from the arid and semi-arid regions of Southern Zacatecas, Mexico. In particular, Stenocereus spp. are important cacti in the region due to its nutritional properties, role as an economic resource, and cultural significance.It is worth emphasising that these cacti traditionally grow wild (i.e., without deliberate cultivation); accordingly, controlled cultivation is uncommon and remains understudied. With the aim of producing a formal, comprehensive analysis and compendium, the data were collected across multiple phenological stages to provide a complete representation of the plant development cycle, from vegetative growth through to fruiting. To achieve this, the collection process combined high-resolution multispectral imaging with field spectrometry in the 400-700 nm range. Standardized acquisition protocols were applied in field conditions to capture consistent reflectance data, and environmental variables such as illumination, temperature, and geographic coordinates were recorded for each session to ensure reproducibility. The dataset integrates several components: (i) multispectral images that provide spatial information on canopy and structural characteristics, (ii) field spectral signatures with detailed reflectance values for each sampled plant, and (iii) metadata describing phenological stage, acquisition date and time, environmental conditions, and equipment settings. For subsequent analysis, data was preprocessed and normalized to enable reliable comparisons between growth stages and across acquisition sessions, resulting in a clean, structured resource ready for computational analysis. In this regard, this dataset has been organized to facilitate its direct application across multiple research and development contexts. Specifically, potential applications include the training and validation of machine learning and computer vision models for automated phenological stage classification, harvest time estimation, and development of species-specific vegetation indices. Moreover, owing to its standardized design, the resource can serve as a benchmark for comparing methods, validating algorithms, and supporting reproducible workflows in precision agriculture and remote sensing. Beyond Stenocereus queretaroensis, the documented acquisition and preprocessing methodology can be replicated or adapted to generate similar multimodal datasets for other climate-resilient crops, particularly those cultivated in arid and semi-arid regions. This could enable comparative analyses across species and provide a reference for extending multimodal sensing approaches to underrepresented plants of ecological and economic importance.

Why it matches plant phenotyping methods植物の生育段階・生理・構造特性を対象に、標準化されたマルチスペクトル画像とフィールド分光データを収集・前処理した再利用可能なデータセットであり、ベンチマークやアルゴリズム検証を目的とするため、フェノタイピング手法が中心です。

abstractThis data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo)
Reproduction assets foundThe paper's own multimodal phenotyping dataset (multispectral/RGB images, spectral signatures, NDVI products, metadata, and example MATLAB scripts) is publicly deposited on Mendeley Data with explicit direct URL and DOI.
Dataset · public) at ∼1750 m a.s.l., under semi-arid temperate conditions with spring temperatures ranging 20–33°C. The data were collected from the Unit Academic of Electrical Engineering Plantel Jalpa. Data accessibility Repository name: Multimodal_Cactaceae_Dataset_25 Data identification number: doi:10.17632/skw8tjc82f.1 Direct URL to data: https://data.mendeley.com/datasets/skw8tjc82f/1 Instructions for accessing these data: click on the direct URL to obtain the multimodal data from Mendeley Dataset Repository. Related research article None 1. Value of the Data • These data provide a unique, non-invasive resource for studying Stenocereus spp. physiology. The integrated collection of high-resolution multOpen asset ↗Mendeley Data · doi:10.17632/skw8tjc82f.1lines:32-58
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published20 Nov 2025Data in briefCited by 4 · OpenAlex ↗

Pomegranate disease detection and classification dataset for deep learning applications: A case study from Halabja city.

Field / plotFruitClassificationStress / disease detectionDisease symptoms / severity

Timely and accurate detection of pomegranate fruit diseases is critical for minimizing crop losses, preserving fruit quality, and supporting sustainable agricultural practices. This study introduces the Halabja Pomegranate Fruit Disease Image Dataset, a systematically compiled collection of images from orchards in one of Iraq's major pomegranate-producing regions. The dataset comprises 2178 original images and 28,314 augmented images, categorized into four specific classes: ectomyelois ceratoniae, colletotrichum spp., sunburn, and healthy fruit samples. To create an ecological setting and ensure significant class variation, images were captured in natural outdoor environments. A standard preprocessing step was applied, which involved resizing all images to 512×512 pixels and using several image augmentation techniques to improve the flexibility and robustness of machine learning models. The unique characteristics of this dataset make it highly suitable for developing machine learning and deep learning models aimed at plant disease detection and other computer vision tasks in precision agriculture. Its contextual relevance and content diversity make it valuable for building an effective diagnostic tool capable of functioning in real field conditions.

Why it matches plant phenotyping methods植物病害状態を画像で分類するデータセットの構築が中心で、再利用可能な植物表現型データとして適格です。

abstractThis study introduces the Halabja Pomegranate Fruit Disease Image Dataset
Reproduction assets foundThe paper is a data descriptor for the authors' own Halabja Pomegranate Fruit Disease Image Dataset (2178 original + 28,314 augmented images), publicly deposited on Zenodo with an explicit direct URL matching an allowed URL. This is a paper-specific public plant-image/phenotyping asset.
Dataset · publicasses: Colletotrichum spp. (anthracnose), Ectomyelois ceratoniae (fruit borer), sunburn, and healthy fruit. Data source location Pomegranate orchards in Halabja city, Kurdistan region, Iraq (location code: 46,018). Data accessibility Repository name: Zenodo Data identification number: 10.5281/zenodo.15856012 Direct URL to data: https://zenodo.org/records/15856012 Halabja Pomegranate Fruit Disease Image Dataset. Zenodo [ 1 ]. Related research article None 1. Value of the Data • Regional Uniqueness: This dataset is the first publicly available collection of pomegranate fruit disease images from Halabja, in the Kurdistan Region of Iraq, an area renowned for its high-quality pomegranate proOpen asset ↗Zenodo · 10.5281/zenodo.15856012lines:1-52
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published20 Nov 2025Sensors (Basel, Switzerland)Cited by 11 · OpenAlex ↗

DLCPD-25: A Large-Scale and Diverse Dataset for Crop Disease and Pest Recognition.

Field / plotClassificationDisease symptoms / severity

The accurate identification of crop pests and diseases is critical for global food security, yet the development of robust deep learning models is hindered by the limitations of existing datasets. To address this gap, we introduce DLCPD-25, a new large-scale, diverse, and publicly available benchmark dataset. We constructed DLCPD-25 by integrating 221,943 images from both online sources and extensive field collections, covering 23 crop types and 203 distinct classes of pests, diseases, and healthy states. A key feature of this dataset is its realistic complexity, including images from uncontrolled field environments and a natural long-tail class distribution, which contrasts with many existing datasets collected under controlled conditions. To validate its utility, we pre-trained several state-of-the-art self-supervised learning models (MAE, SimCLR v2, MoCo v3) on DLCPD-25. The learned representations, evaluated via linear probing, demonstrated strong performance, with the SimCLR v2 framework achieving a top accuracy of 72.1% and an F1 score (Macro F1) of 71.3% on a downstream classification task. Our results confirm that DLCPD-25 provides a valuable and challenging resource that can effectively support the training of generalizable models, paving the way for the development of comprehensive, real-world agricultural diagnostic systems.

Why it matches plant phenotyping methods作物の病害・健全状態を画像で認識する大規模公開ベンチマークデータセットを構築・評価しており、植物状態の画像ベース表現型解析基盤が中心です。害虫認識も含まれますが、病害・健全状態の評価は植物フェノタイピングに該当します。

abstractwe introduce DLCPD-25, a new large-scale, diverse, and publicly available benchmark dataset.
Reproduction assets foundThe paper introduces DLCPD-25, a public crop pest/disease image dataset (221,943 images, 203 classes), with an explicit Data Availability Statement pointing to the authors' GitHub repository containing all image data and documentation.
Dataset · publicThe DLCPD-25 dataset introduced and analyzed in this study is publicly available at: https://github.com/hwzhanng/DLCPD-25-Dataset (accessed on 20 October 2025). The repository provides access to all image data, and relevant documentation used in this research.Open asset ↗https://github.com/hwzhanng/DLCPD-25-Dataset · DLCPD-25lines:141-207
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published10 Nov 2025Scientific DataCited by 3 · OpenAlex ↗

Annotated 3D Point Cloud Dataset of Broad-Leaf Legumes Captured by High-Throughput Phenotyping Platform.

Common beanCowpeaLiDAR / point cloudMultispectral / hyperspectralLeafStem / branchWhole plant / canopy / plot / fieldAnnotation / quality controlCalibration / preprocessingSegmentation

This data descriptor presents novel, annotated 3D point cloud plant scans generated by a high-throughput phenotyping platform (LeasyScan, ICRISAT, India). It focuses on broad-leaf legume species (mungbean, common bean, cowpea, and lima bean). The dataset, generated by PlantEye(R) F600 technology, captures multispectral 3D scans of plant canopies. It includes 223 scans, providing detailed organ-level segmentation annotations for embryonic leaves, leaves, petioles, stems, and whole plants. The dataset fills a critical gap in plant phenomics research by offering a base of annotated data to support AI model development efforts in 3D computer vision. Data preprocessing, annotation procedures, and potential applications in crop research disciplines are further discussed. The dataset, preprocessing code, annotations, and a MIAPPE-compliant data sheet are also presented via the GitHub repository for further updates and expansion.

Why it matches plant phenotyping methods植物フェノタイピングプラットフォームで取得した3D点群と器官レベル注釈を提供するデータセットで、再利用可能な画像解析・AI開発基盤が中心です。

abstractThis data descriptor presents novel, annotated 3D point cloud plant scans generated by a high-throughput phenotyping platform (LeasyScan, ICRISAT, India).
Reproduction assets foundThe paper's own annotated 3D point cloud dataset (223 scans of legumes with organ-level segmentation annotations), raw scanner data, MIAPPE metadata, and preprocessing/cuboid-generation/baseline-evaluation code are publicly deposited on Figshare and mirrored on GitHub.
Code · publicinto this software. All the code and data are also available as the GitHub (https://github.com/kit-pef-czu-czOpen asset ↗GitHubpdf-page:2 lines:1-58
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published7 Nov 2025Plant methodsCited by 4 · OpenAlex ↗

Hyperspectral image analysis for classification of multiple infections in wheat.

WheatMultispectral / hyperspectralLeafClassificationStress / disease detectionDisease symptoms / severity

Plant diseases can cause heavy yield losses in arable crops resulting in major economic losses. Effective early disease recognition is paramount for modern large-scale farming. Since plants can be infected with multiple concurrent pathogens, it is important to be able to distinguish and identify each disease to ensure appropriate treatments can be applied. Hyperspectral imaging is a state-of-the art computer vision approach, which can improve plant disease classification, by capturing a wide range of wavelengths before symptoms become visible to the naked eye. Whilst a lot of work has been done applying the technique to identifying single infections, to our knowledge, it has not been used to analyse multiple concurrent infections which presents both practical and scientific challenges. In this study, we investigated three wheat pathogens (yellow rust, mildew and Septoria), cultivating co-occurring infections, resulting in a dataset of 1447 hyperspectral images of single and double infections on wheat leaves. We used this dataset to train four disease classification algorithms (based on four neural network architectures: Inception and EfficientNet with either a 2D or 3D convolutional layer input). The highest accuracy was achieved by EfficientNet with a 2D convolution input with 81% overall classification accuracy, including a 72% accuracy for detecting a combined infection of yellow rust and mildew. Moreover, we found that hyperspectral signatures of a pathogen depended on whether another pathogen was present, raising interesting questions about co-existence of several pathogens on one plant host. Our work demonstrates that the application of hyperspectral imaging and deep learning is promising for classification of multiple infections in wheat, even with a relatively small training dataset, and opens opportunities for further research in this area. However, the limited number of Septoria and yellow rust + Septoria samples highlights the need for larger, more balanced datasets in future studies to further validate and extend our findings under field conditions.

Why it matches plant phenotyping methods小麦葉の感染状態をハイパースペクトル画像と深層学習で分類する手法が研究の中心であり、植物病害状態の表現型推定に該当する。

abstractHyperspectral imaging is a state-of-the art computer vision approach, which can improve plant disease classification
Reproduction assets foundThe paper's Data availability statement explicitly deposits the authors' training/deployment/testing code for the hyperspectral wheat disease classification models in a public GitHub repository. The 1447-image hyperspectral dataset itself has no stated public deposit, so it is not included as an asset.
Code · publicCode for training, deploying and testing the models can be found at https://github.com/mc2295/hyperspectralplants .Open asset ↗mc2295/hyperspectralplantslines:140-218
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Nov 2025Data in briefCited by 1 · OpenAlex ↗

Image dataset of ten durian diseases captured in real-field conditions from a family orchard in Vinh Long, Vietnam.

Field / plotFlowerLeafRootStem / branchClassificationDisease symptoms / severity

This dataset comprises 5452 images of durian plant parts-including leaves, flowers, branches, stems, and roots-affected by ten common disease classes. The images were captured from one family-owned durian orchard and four nearby orchards in Vinh Long Province, Vietnam. Each class contains approximately 405-427 raw images, photographed using an iPhone 14 under natural field conditions. These conditions simulate typical farmer photography practices, featuring varied angles, inconsistent lighting, and complex environmental backgrounds, resulting in significant visual noise. All raw JPEG images were manually reviewed and cropped on macOS systems using MacBook devices equipped with Apple M4 chips to focus on disease-affected regions, reduce file size, and minimize background noise. The processed, cropped images are provided in PNG format with variable dimensions. Images were resized to 224×224 pixels only during model training for machine learning experiments. Disease symptoms were verified in collaboration with plant pathologists to ensure accurate classification. This dataset is publicly available on Mendeley Data and is suitable for developing and evaluating machine learning models in plant disease classification. It is particularly valuable for testing model performance under real-world, noisy conditions and for supporting the creation of mobile or edge-based diagnostic tools in agriculture.

Why it matches plant phenotyping methods植物病徴を画像で直接捉えた公開データセットで、植物病害状態の分類モデル開発・評価を主目的とするため、表現型計測データセットとして中心的です。

abstractThis dataset comprises 5452 images of durian plant parts-including leaves, flowers, branches, stems, and roots-affected by ten common disease classes.
Reproduction assets foundThe paper is a Data in Brief describing a public durian disease image dataset (5452 field images, ten classes) deposited on Mendeley Data with an explicit DOI and direct URL, matching an allowed URL. This is a paper-specific, publicly available image dataset directly reproducing the paper's phenotyping measurements. No
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/mhjwyb5p48 Direct URL to data: https://data.mendeley.com/datasets/mhjwyb5p48/1Open asset ↗Mendeley Data · 10.17632/mhjwyb5p48lines:47-125
Code / dataset availability confirmedOpenAlex · checked 13 Sept 2026
Published4 Nov 2025Earth system science dataCited by 2 · OpenAlex ↗

Countrywide digital surface models and vegetation height models from historical aerial images

Aerial / UAVPhotogrammetry / SfM / MVSStereo2D/3D reconstructionPlant / canopy height

Abstract. Historical aerial images, captured by film cameras in the previous century, are valuable resources for quantifying Earth's surface and landscape changes over time. In the post-war period, these images were often acquired to create topographic maps, resulting in the acquisition of large-scale aerial photographs with stereo coverage. Photogrammetric techniques applied to these stereo images enable the extraction of 3D information to reconstruct digital surface models (DSMs) and orthoimages. Here, we present a highly automated photogrammetric approach for generating countrywide DSMs of Switzerland, at a 1 m resolution, from approximately 32 000 scanned aerial stereo images acquired between 1979 and 2006, with known exterior and interior orientation. We derived four countrywide DSMs for the epochs 1979–1985, 1985–1991, 1991–1998, and 1998–2006. From the DSMs, we generated corresponding countrywide vegetation height models (VHMs). We assessed the quality of the historical DSMs at the country scale and within six representative study sites, evaluating the vertical accuracy and the completeness of image matching across different land cover types. Mean completeness ranged from 64 % for “glacial and perpetual snow” to 98 % for “sealed surfaces”, with a value of 93 % for the “closed forest” class. Across Switzerland, the median elevation accuracy of the historical DSMs compared with a reference digital terrain model (DTM) on sealed surface points ranged from 0.08 to 0.16 m, with a normalized median absolute deviation (NMAD) of around 0.8 m and a maximum root mean square error (RMSE) of 1.20 m. Similar accuracies are obtained when comparing historical DSMs with measured geodetic points. The VHMs generated in this study enabled the detection of major changes in forest areas due to windstorm damage, forest dynamics, and growth. This work demonstrates the feasibility of generating accurate, very-high-resolution DSM time series (spanning three decades) and VHMs from historical aerial images of the entire surface of Switzerland in a highly automated manner. The VHMs are already being used to estimate countrywide biomass changes. The countrywide DSMs and VHMs for the four epochs, along with auxiliary data, are available online at https://doi.org/10.16904/envidat.528 (Marty et al., 2024) and can be used to quantify long-term elevation changes and related processes across different surfaces.

Why it matches plant phenotyping methods歴史的航空画像から植生高モデルを生成する自動写真測量法を開発・精度評価し、森林の高さ変化という植物キャノピー形質を抽出しているため、測定法が中心的である。

abstractFrom the DSMs, we generated corresponding countrywide vegetation height models (VHMs).
Reproduction assets foundThe paper's countrywide DSMs, VHMs, and auxiliary rasters (matching mask, vegetation mask, metadata shapefile) for four epochs are deposited publicly on EnviDat with an explicit DOI. These vegetation height models are the paper's plant/canopy phenotyping measurements. No author analysis code or trained models are named
Dataset · publicDatasets can be accessed from EnviDat ( https://doi.org/10.16904/envidat.528 , Marty et al., 2024). The following files are available for the four epochs: countrywide digital surface model (DSM), hillshaded DSM, and vegetation height models (VHMs).Open asset ↗Envidat · 10.16904/envidat.528lines:249-256
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 Nov 2025Data in briefCited by 1 · OpenAlex ↗

WMC-Leafset: A dataset of wax gourd and Mangalore cucumber plants for leaf miner and pest infestation diseased object detection.

MelonField / plotLeafClassificationObject detectionDisease symptoms / severity

Wax gourd ( Benincasa hispida (Thunb.) Cogn.) and Mangalore Cucumber (Cucumis melo L. subsp. agrestis var. conomon) are nutritionally rich, mineral-dense crops with a short growing cycle, making them a preferred choice for cultivation among farmers across the country. The Mangalore cucumber, also known as the culinary cucumber, Indian yellow cucumber, or Japanese pickling melon, is widely used in Asian cuisine for pickling. While proper nutrient management is essential for optimal growth, disease control poses a significant challenge in ensuring healthy yields, as disease can rapidly spread from one leaf to another, affecting larger areas of the field and reducing crop yield. Since cucurbits grow close to the soil, they spread across the ground, exhibit dense canopies, and often overlap with neighboring plants. Early detection is crucial to ensure sustainable cultivation, food security, and increased crop productivity. To address this challenge, we collected a dataset comprising 3200 images that includes image samples of Wax gourd and Mangalore cucumber plants affected by leaf miner, pests and image samples of healthy leaves. The Cucurbitaceae datasets that are available in the public domain lack representation of the Mangalore cucumber and Wax gourd varieties. To the best of our knowledge, no publicly available dataset exists for the Wax gourd. Moreover, existing datasets typically contain images captured under controlled greenhouse conditions with plain backgrounds, featuring a single leaf per image. They exhibit low background complexity and limit the scope to detect diseases at the object level, including multiple diseases present on a single leaf or plant. The uniqueness of the proposed dataset lies in addressing this gap by providing field-level images of cucurbits. These images capture variations in soil, overlapped leaves, complex background, varying angles and distances, weeds, and human interference. This makes the dataset suitable for training object detection models capable of identifying single and multiple disease instances, and it can also be effectively used for classification tasks to distinguish between healthy and diseased leaves. It supports advancement in deep learning, feature extraction, segmentation and pattern recognition tasks. Additionally, the dataset serves as a valuable resource for plant pathologists, agronomists and agricultural experts in disease detection, monitoring and management, thereby promoting sustainable agricultural practices. By offering open access, this dataset promotes collaboration within the scientific community to facilitate the development of robust disease detection, identification, and disease control, thus enhancing farming practices and increasing agricultural yields and advancing food security.

Why it matches plant phenotyping methods植物の病害状態を画像から検出・分類する公開データセットが研究の中心であり、植物病害の画像ベース表現型解析に該当する。

abstractwe collected a dataset comprising 3200 images that includes image samples of Wax gourd and Mangalore cucumber plants affected by leaf miner, pests and image samples of healthy leaves.
Reproduction assets foundThe paper is a Data in Brief article describing the WMC-Leafset dataset of 3200 annotated field images of wax gourd and Mangalore cucumber plants for leaf miner/pest object detection. The dataset is publicly deposited on Mendeley Data with an explicit direct URL and DOI, making it a paper-specific, publicly actionable,
Dataset · public6° 44′ 46″ east, latitude of 12.16999° or 12° 10′ 12″ north and Mangalore cucumber images were collected from Hulimahu village in longitude of 76.73329° or 76° 43′ 60″ east, latitude 12.1598° or 12° 9′ 35″ north Data accessibility Repository name: WMC-Leafset Data identification number: 10.17632/8m2ytxd4dg.4 Direct URL to data: https://data.mendeley.com/datasets/8m2ytxd4dg/4 Related research article [ 11 ] M. A. Keerthi Prasad, N. Shobha Rani, M. A. Sangamesha and K. V. Vinay, ``Identification and Detection of Leaf Miner, Pest Infestation in Cucurbitaceae Family in Real-Time Infield Scenarios using YOLOv5s Object Detection Model,'' 2024 11th International Conference on Computing for SustainaOpen asset ↗10.17632/8m2ytxd4dg.4lines:35-65
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 Nov 2025Data in briefCited by 1 · OpenAlex ↗

Leaf functional trait dataset of 93 dominant woody species from the central Western Ghats, India.

Field / plotLeafMorphology / geometry measurementLeaf traits

We present a comprehensive dataset of qualitative and quantitative leaf functional traits for 93 dominant woody species representing two distinct leafing phenologies and three growth forms from the central Western Ghats of India. Quantitative assessments were conducted for nine key traits: leaf area (LA), mean thickness (LTH), specific leaf area (SLA), leaf dry matter content (LDMC), leaf tissue density (LTD), leaf nitrogen concentration (Leaf N), carbon-to-nitrogen ratio (C/N), and phytolith yield, following standard protocols. For each species, 30 leaves were sampled from a minimum of five individuals, totalling 2790 leaf samples. Qualitative traits, including leaf shape, margin, surface, texture, apex, base, type, and latex presence, were recorded in the field and validated using field manuals. The majority of species sampled were evergreen (74 %), with deciduous species comprising the remainder. Given the growing importance of plant functional traits in ecological research, this dataset offers valuable species-level leaf trait information at the regional scale. The phytolith yield data, in particular, represent one of the few globally available datasets, providing essential baselines for palaeoecological research and enabling quantitative reconstruction of vegetation composition and environmental change over millennial timescales.

Why it matches plant phenotyping methods植物の葉形質を標準化プロトコルで体系的に収集した再利用可能なデータセット論文であり、データセット自体が中心的な成果である。

abstractWe present a comprehensive dataset of qualitative and quantitative leaf functional traits for 93 dominant woody species
Reproduction assets foundThe paper's own leaf functional trait dataset (2790 leaves, 93 woody species, central Western Ghats) is publicly deposited on Zenodo with an explicit DOI/URL given in the article.
Dataset · publicduals per species. Quantitative leaf functional traits were analyzed following the standard protocol [ 1 , 2 ]. Data source location Country: India Sampling site: Gerusoppa Reserve Forest, Central Western Ghats (14°12′ N to 14°24′ N and 74°36′ E to 74°48′ E) Data accessibility Repository name: Zenodo Data identification number: https://doi.org/10.5281/zenodo.16717435 Direct URL to data: https://doi.org/10.5281/zenodo.16717435 Related research article None 1 Value of the Data • This dataset provides high-resolution leaf-level data ( n = 2790) on 17 functional traits for 93 dominant woody species of the central Western Ghats, supporting trait-based ecological research. • It enables assessmentOpen asset ↗Zenodo · 10.5281/zenodo.16717435lines:1-54
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published1 Nov 2025Data in briefCited by 2 · OpenAlex ↗

A comprehensive dataset of agarwood tree ( Aquilaria Malaccensis ) leaf images for disease analysis in Brunei Darussalam.

Field / plotLeafClassificationDisease symptoms / severity

The visual diagnosis based on foliar traits remains a cornerstone technique for the early identification of biotic stress, for instance, disease and pest infestations, in many economically valuable crops, including Aquilaria Malaccensis (agarwood). As a species of immense commercial and ecological significance, Aquilaria Malaccensis is particularly vulnerable to a range of pathogens and insect threats that can severely compromise resin production and tree viability. With the increasing integration of disruptive sustainable agricultural technologies, such as artificial intelligence (AI), especially in plant phenotyping and pathology, the development of robust and generalizable AI models hinges on the availability of large-scale and high-resolution image datasets. However, the current lack of such curated datasets for agarwood poses a substantial bottleneck to progress in automated identification systems. This deficiency limits the ability of scientists, technologists, and plant health experts to leverage machine learning and computer vision techniques for timely, accurate, and scalable solutions to different stresses in agarwood disease and pest management, including nematodes, viroids, viruses, pests, phytoplasmas, bacteria, fungi, and Protozoa. This paper presents a dataset of pests and diseases affecting agarwood trees, which impact farmers. It includes a total of 5472 leaf images classified into 14 categories. These categories consist of 8 types of agarwood diseases, 5 types of pests, and 1 category of healthy leaf images, encompassing both insect-damaged and healthy leaves. The images were captured using a PowerShot G7X Mark III camera. The images were captured from three different agarwood plantation sites of Batong, Benutan, and Bukit Silat in 2024, led by the Institute for Biodiversity and Environmental Research (IBER), Universiti Brunei Darussalam, by Botanical Research Centre (UBD BRC) scientists and biologists. This dataset is particularly valuable for training and validating deep learning (DL), computer vision, and machine learning algorithms aimed at identifying agarwood diseases and pests in agarwood leaves. Offering researchers and learners a robust data resource for analyzing and improving agarwood plant health through the development of advanced computational models. The designed models are vital and hold immense practical value for farmers, equipping them with the tools that timely detect and identify diseases in their agarwood trees, empowering them to make informed decisions and potentially intensify their profits.

Why it matches plant phenotyping methods葉画像から病害・害虫による植物状態を識別するための大規模データセットを構築しており、データ取得と再利用可能な解析基盤が研究の中心であるため。

abstractThis dataset is particularly valuable for training and validating deep learning (DL), computer vision, and machine learning algorithms aimed at identifying agarwood diseases and pests in agarwood leaves.
Reproduction assets foundThe paper is a data descriptor for a public agarwood leaf image dataset (5472 images, 14 classes) deposited on Zenodo and Mendeley, with explicit direct URLs and DOIs provided in the Data Accessibility section. This is the paper's own phenotyping image dataset, publicly available and actionable. No separate author code
Dataset · publicc.iber.ubd.edu.bn ), Universiti Brunei Darussalam, Gadong, BE1410, Brunei Darussalam Data accessibility Repository name: Zendo and Mendeley Repository Title: Agarwood Leaf Image Dataset for Pest and Disease Analysis in Real-World Environment Data identification number: https://doi.org/10.5281/zenodo.14842099 Direct URL to data: https://zenodo.org/records/14842100 Direct URL to data: https://data.mendeley.com/datasets/8f8wtr9zwn/2 Related research article Shafik, W., Tufail, A., De Silva, L.C. et al. A lightweight deep learning model for multi-plant biotic stress classification and detection for sustainable agriculture. Sci Rep 15, 12,195 (2025). https://doi.org/10.1038/s41598-025-90487-Open asset ↗Zenodo · 10.5281/zenodo.14842099lines:36-67
Dataset · publicg, BE1410, Brunei Darussalam Data accessibility Repository name: Zendo and Mendeley Repository Title: Agarwood Leaf Image Dataset for Pest and Disease Analysis in Real-World Environment Data identification number: https://doi.org/10.5281/zenodo.14842099 Direct URL to data: https://zenodo.org/records/14842100 Direct URL to data: https://data.mendeley.com/datasets/8f8wtr9zwn/2 Related research article Shafik, W., Tufail, A., De Silva, L.C. et al. A lightweight deep learning model for multi-plant biotic stress classification and detection for sustainable agriculture. Sci Rep 15, 12,195 (2025). https://doi.org/10.1038/s41598-025-90487-1 . 1. Value of the Data • The dataset comprises 5472 high-quOpen asset ↗Mendeley · 8f8wtr9zwnlines:36-67
Code / dataset availability confirmedEurope PMC · bioRxiv · Crossref · checked 14 Sept 2026
Published23 Oct 2025bioRxivCited by 1 · OpenAlex ↗

Benchmarking remote sensing methods to capture plant functional diversity from space

Field / plotChlorophyll fluorescenceMultispectral / hyperspectralThermalLeafWhole plant / canopy / plot / fieldGrowth / time-series analysisLeaf traitsPhotosynthesis / fluorescenceYield / yield components

ABSTRACT The development of remote sensing methods to estimate plant functional diversity is limited by mismatches between ecology and remote sensing sampling schemes, and the limited representativeness of local field campaigns. The Biodiversity Observing System Simulation Experiment (BOSSE) provides a modeling framework for benchmarking new methodologies. We used BOSSE to simulate 180 different synthetic “Scenes” encompassing a two-year-long time series of plant trait maps and imagery of hyperspectral reflectance factors, spectral indices, sun-induced chlorophyll fluorescence, land surface temperature, and estimates of plant traits (optical traits). We used these simulations to answer five fundamental, yet unsolved, questions: Q1. How should remote sensing characterize functional diversity in large surfaces (sites)? Diversity metric values saturate with the number of pixels involved, hampering comparisons between plant traits and remote sensing estimates in large areas. The average value of metrics computed over small samples should be used instead. Q2. Which sources of spectral information (or combinations thereof) can best capture plant functional diversity at the site scale? Accounting for background effects is the key. Optical traits (remote sensing estimates of plant traits) are the best estimators for plant functional diversity. Other variables succeed when filtered out of the soil pixels; their combination did not yield additional advantages. Q3. How should remote sensing estimates be validated/compared with plant functional diversity measurements? Leaf area index (LAI) is a better proxy of abundance than the pixel for Q Rao, but not for variance-based partitioning. It is more sensitive to sample size, but also more resistant to suboptimal spatial resolution. Q4. When (in the phenological year) can remote sensing best capture site-scale plant functional diversity? The estimation error decreased with LAI and stabilized at values above 1 m²/m². Q5. Which approaches and remote sensing variables are more resistant to the effects of suboptimal spatial resolution? Optical traits, fluorescence, and reflectance factors were the most robust variables. Still, field data resolution needs to be degraded to match the sensor’s resolution. We found a relative spatial resolution threshold of ∼30 % (where the pixel is around three times larger than the plants). Simulation frameworks like BOSSE enable testing methodologies beyond local contexts and address the current shortage of suitable global datasets, supporting the application and development of methods for assessing plant functional diversity with remote sensing. In the future, BOSSE could contribute to understanding observational results, refining and pre-testing new methodologies, and supporting the development of comparable experimental datasets.

Why it matches plant phenotyping methodsBOSSEを用いてリモートセンシングによる植物形質・機能多様性推定手法をシミュレーションベンチマークし、検証・比較する研究であり、植物フェノタイピング手法が中心である。

abstractThe Biodiversity Observing System Simulation Experiment (BOSSE) provides a modeling framework for benchmarking new methodologies.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicvariables, we used the “pyGNDiv” package (https://github.com/JavierPachecoLabrador/pyGNDiv-Open asset ↗JavierPachecoLabrador/pyGNDiv- · pyGNDivpdf-page:11 lines:1-60
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published20 Oct 2025BiogeosciencesCited by 1 · OpenAlex ↗

Isotope discrimination of carbonyl sulfide ( 34 S) and carbon dioxide ( 13 C, 18 O) during plant uptake in flow-through chamber experiments

SunflowerLaboratory / benchtopLeafWhole plant / canopy / plot / fieldPhysiological trait estimationPhotosynthesis / fluorescenceWater status / transpiration

Abstract. Carbonyl sulfide (COS) has been proposed as a proxy for gross primary production (GPP), as it is taken up by plants through a pathway comparable to that of CO2. COS diffuses into the leaf, where it undergoes an essentially one-way reaction in the mesophyll cells, irreversibly catalyzed by the enzyme carbonic anhydrase (CA), and is likely not respired by the leaf. In order to use COS as a proxy for GPP, the mechanisms of COS uptake and its coupling to photosynthesis need to be well understood. Characterizing the isotopic discrimination of COS during plant uptake could provide valuable information on the physiological COS uptake process and may help to constrain the COS budget. This study presents joint measurements of isotope discrimination during plant uptake for COS (CO34S) and CO2 (13CO2 and C18O16O). A C3 plant, sunflower (Helianthus annuus), and a C4 plant, papyrus (Cyperus papyrus), were enclosed in a flow-through plant chamber and exposed to varying light levels. The incoming and outgoing gas compositions were measured online, and discrete air samples were taken for isotope analysis. Simultaneously measuring fluxes and isotope discrimination of both COS and CO2 yielded a unique dataset that includes information on the plant's behavior and allowed for the estimation of stomatal- and mesophyll conductances. The average COS uptake fluxes were 73.3 ± 1.5 pmol m−2 s−1 for sunflower and 107.3 ± 1.5 pmol m−2 s−1 for papyrus (PAR > 0) and displayed virtually no trend with increasing PAR from 200 to 600 µmol m−2 s−1. The mean observed 34Δ for COS was 3.4 ± 1.0 ‰ for sunflower and 2.6 ± 1.0 ‰ for papyrus. 34Δ was stable across all light intensities, which could be explained by a sufficient stomatal opening and low variability in the ratio of mesophyll vs. ambient COS mole fraction, CmS/CaS. For both C3 and C4 plants, for CO2, a negative relationship was observed between the uptake flux and the isotopic discriminations 13Δ and 18Δ. The CO2 uptake and 13CO2 and C16O18O discriminations of sunflower have expected values for a C3 plant, while the low CO2 flux and high 13Δ and 18Δ values observed for papyrus were not in the typical C4 range, which was perhaps due to the relatively low light conditions during our experiments.

Why it matches plant phenotyping methods植物のCOS・CO2取り込み、同位体識別、気孔・葉肉コンダクタンスをフロースルー植物チャンバーで定量する生理的表現型測定が研究の中心であり、再利用可能な測定データセットと推定手法を提示している。

abstractThis study presents joint measurements of isotope discrimination during plant uptake for COS (CO34S) and CO2 (13CO2 and C18O16O).
Reproduction assets foundThe paper's isotope discrimination and gas-exchange dataset from the flow-through chamber experiments is publicly deposited on Zenodo by the authors.
Dataset · publicynthetically available radiation at the top of the chamber, 34 Δ is the discrimination against CO 34 S and LRU is the leaf relative uptake ratio. * n =1 , error states is the single measurement precision instead of the repeatability precision. Download Print Version | Download XLSX Data availability The dataset is available at: https://doi.org/10.5281/zenodo.14677494 (Baartman et al., 2025). Author contributions Conceptualization: SLB, MCK, MEP, LW. Data curation: SLB. Formal analysis: SLB, NUL. Funding acquisition: MCK. Investigation: SLB, SMD, MW, LMJK, LM, AC, SH. Methodology: SLB, SMD, MW, LMJK, MEP. Resources: SMD, MW, LM, SH. Supervision: MEP, TR, MCK. Visualization: SLB, NUL. WritingOpen asset ↗Zenodo · 10.5281/zenodo.14677494lines:652-942
Code / dataset availability confirmedarXiv · OpenAlex · checked 15 Sept 2026
Published19 Oct 2025arXivCited by 0 · OpenAlex ↗

An RGB-D Image Dataset for Lychee Detection and Maturity Classification for Robotic Harvesting

Field / plotRGB-D / ToFFruitClassificationObject detectionFruit / seed / panicle traits

Lychee is a high-value subtropical fruit. The adoption of vision-based harvesting robots can significantly improve productivity while reduce reliance on labor. High-quality data are essential for developing such harvesting robots. However, there are currently no consistently and comprehensively annotated open-source lychee datasets featuring fruits in natural growing environments. To address this, we constructed a dataset to facilitate lychee detection and maturity classification. Color (RGB) images were acquired under diverse weather conditions, and at different times of the day, across multiple lychee varieties, such as Nuomici, Feizixiao, Heiye, and Huaizhi. The dataset encompasses three different ripeness stages and contains 11,414 images, consisting of 878 raw RGB images, 8,780 augmented RGB images, and 1,756 depth images. The images are annotated with 9,658 pairs of lables for lychee detection and maturity classification. To improve annotation consistency, three individuals independently labeled the data, and their results were then aggregated and verified by a fourth reviewer. Detailed statistical analyses were done to examine the dataset. Finally, we performed experiments using three representative deep learning models to evaluate the dataset. It is publicly available for academic

Why it matches plant phenotyping methodsライチ果実の成熟段階という植物器官の状態をRGB-D画像から分類するデータセットを構築し、アノテーション検証と深層学習モデル評価を行っており、表現型取得・評価手法が中心である。

abstractwe constructed a dataset to facilitate lychee detection and maturity classification.
Reproduction assets foundThe authors publicly release the paper's lychee RGB-D image dataset (raw/augmented RGB images, depth maps, detection and maturity annotations) and the Python scripts for data augmentation, image similarity comparison, and annotation in the same GitHub repository.
Dataset · publicchees, the non-augmented models produced misclassifications with lower recognition and accuracy, whereas the augmented models avoided these issues. Overall, the results demonstrate that the data augmentation method effectively improves the comprehensive performance of the models. 5. Data Availability The dataset is available at:https://github.com/SeiriosLab/Lychee. The Python scripts for data augmentation, image similarity comparison, and annotation are available within the same repository under the tree/main/script directory.Open asset ↗SeiriosLab/Lycheepdf-raw-page:13 lines:1-55
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published16 Oct 2025Data in briefCited by 0 · OpenAlex ↗

Smartphone image dataset for turmeric plant leaf disease from Bangladesh spice fields.

Field / plotRGB / grayscaleLeafClassificationDisease symptoms / severity

Agriculture is key to sustaining life and economic development, and crops like turmeric are essential for everyday application and economic viability. Turmeric crops are very prone to foliar disease, which has a great impact on yield and quality. Early detection of the diseases is of great significance to farming practitioners since manual observation is generally time-consuming and unreliable. To surpass this challenge, a comprehensive dataset has been developed to facilitate the generation of an automatic disease recognition system. The dataset comprises 865 images of original turmeric leaves and 3496 images of augmented turmeric leaves, both infected and healthy, with four classes of diseases: aphid attack, blotch, leaf spot, and healthy leaves. All the leaves were captured from different angles to offer variability and clarity, with particular emphasis on high-quality and diversified data. Through this dataset, a precise and efficient identification process can be realized, which will aid agriculture practitioners in recognizing diseases at an early stage and reducing crop losses. This paper seeks to improve agricultural productivity, crop quality, and the overall growth and sustainability of the agricultural economy using state-of-the-art deep learning models, such as EfficientNetB7 and ResNet152, for precise and interpretable disease classification. The proposed approach achieves high accuracy, with EfficientNetB7 attaining 98.67 % and ResNet152 reaching 97.87 %. Additionally, this research lays the groundwork for scalable and affordable disease detection technology, allowing agricultural practitioners to maximize crop yield and achieve long-term food security using smart tools.

Why it matches plant phenotyping methodsターメリック葉の病害状態を画像から分類するデータセットと深層学習手法が研究の中心であり、植物病害フェノタイピングに該当する。

abstracta comprehensive dataset has been developed to facilitate the generation of an automatic disease recognition system.
Reproduction assets foundThe paper is a Data in Brief article describing a turmeric leaf disease image dataset (865 original and 3496 augmented smartphone images) collected by the authors, with the dataset publicly deposited on Mendeley Data. This is a paper-specific, publicly available plant image/phenotyping asset with a direct URL matching,
Dataset · publicEkdonto village turmeric field in Pabna (latitude: 24.071123635779465, longitude: 89.34558471048882) 4. Tebunia village turmeric field in Pabna (latitude: 24.070634252228754, longitude: 89.20332505175169) Data accessibility Repository name: Mendeley Data Data identification number: DOI: 10.17632/jtttfbx342.1 Direct URL to data: https://data.mendeley.com/datasets/jtttfbx342/1 1. Value of the Data • This dataset generates a wealth of visual information on leaf diseases of turmeric, which is a good resource to train machine learning models. The models can be constructed to differentiate well between healthy and diseased leaves so that the diseases can be diagnosed early and accurately in agricuOpen asset ↗Mendeley Data · 10.17632/jtttfbx342.1lines:1-47
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published9 Oct 2025Data in briefCited by 0 · OpenAlex ↗

Central India Medicinal Plant Dataset (CIMPD).

Field / plotLeafDisease symptoms / severity

In the present scenario, medicinal plants play a crucial role in promoting a healthy lifestyle by protecting against numerous diseases. They also hold significant potential as a source of income, particularly for rural populations across the globe. Plants used for herbal medicine are known as medicinal plants, and each part of these plants may be utilized for medicinal purposes. Further, medicinal plants are beneficial in enhancing the human immune system. In this research, a new medicinal plant named as Central India Medicinal Plant Dataset (CIMPD) has been developed to support significant research in human health. The dataset contains 9130 leaf images (both healthy and unhealthy) from 23 medicinal plant species. These images were collected from various locations in central India. The entire work was carried out over a period of five months, which included plant selection, leaf collection, image capturing, and data organization into folders. This dataset provides comprehensive information, including the botanical name, common name, geographical origin, healthy and unhealthy leaf images, and medicinal uses of the plants. It serves as a valuable resource for research in machine learning, computer vision, and related domains. Additionally, it will enable the development and evaluation of methodologies for disease detection, plant identification, and other relevant applications.

Why it matches plant phenotyping methods健康・不健康な葉画像を含む再利用可能なデータセットを構築し、植物の病害状態を画像から判定する研究基盤として提供しているため、画像ベースの植物状態計測に該当する。植物同定も含むが、データセット構築自体が中心である。

abstractThe dataset contains 9130 leaf images (both healthy and unhealthy) from 23 medicinal plant species.
Reproduction assets foundThe paper is a data descriptor for the Central India Medicinal Plant Dataset (CIMPD), a public Kaggle dataset of 9130 healthy/unhealthy medicinal plant leaf images from 23 species, directly reproducing the paper's phenotyping (leaf image) measurements. The ResNet18 feature-visualization analysis code is not explicitly,
Dataset · publichas 9130 images from 23 classes. Within the dataset, there’s an unequal distribution of samples among various classes. Data source location For this project, a large no of gardens of various places of central India has visited to collect the medicinal plant leaves. Data accessibility Repository name: Kaggle Direct URL to data: https://www.kaggle.com/datasets/satyamtomar08/indian-medicinal-plant-dataset 1. Value of the Data • The development of a medicinal plant dataset plays a crucial role in the exploration of advanced machine learning models for significant investigations such as plant identification, disease detection, crop management, and more [ [1] , [2] , [3] , [4] ]. • This plant leafOpen asset ↗Kagglelines:1-54
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published8 Oct 2025Data in briefCited by 1 · OpenAlex ↗

A comprehensive annotated image dataset for deep learning analysis of eggplant leaf diseases.

Eggplant / aubergineField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

The Eggplant Leaf Disease Dataset was meticulously developed to address challenges in accurately identifying diseases that threaten eggplant crops, a vital agricultural resource worldwide. This dataset includes 3116 high-resolution images captured between March and May 2024 from two major agricultural regions in Bangladesh, representing real-world conditions. It comprises 10 distinct disease classes-Aphids, Cercospora Leaf Spot, Defect Eggplant, Flea Beetles, Fresh Eggplant, Fresh Eggplant Leaf, Leaf Wilt, Phytophthora Blight, Powdery Mildew, and Tobacco Mosaic Virus-making it the most comprehensive dataset for eggplant diseases to date. To enhance its utility, rigorous data augmentation techniques, including flipping, rotating, shearing, shifting, noise addition, and brightness adjustment, were applied. This expanded the dataset to 10,000 images, ensuring its robustness for machine learning applications. Expert annotations further enhance its quality, providing critical insights for precise disease classification. Our Proposed CBAM-EfficientNetB0 model had an amazing classification accuracy of 98.70 %, which was much better than the baseline architectures. ResNet50 only got 32.60 %, VGG16 got 73.00 %, and VGG19 got 68.00 %. The proposed model's better performance shows that combining channel and spatial attention through CBAM with EfficientNetB0's feature extraction abilities works well. This architecture does a good job of picking out the distinguishing features in eggplant leaf images, which makes it possible to accurately identify diseases. The dataset and model work together to make AI-powered early disease detection, automated monitoring, and decision support in precision agriculture possible. These tools help farmers use sustainable farming methods by making timely interventions, reducing the need for manual inspection, and increasing crop productivity and food security.

Why it matches plant phenotyping methodsナス葉の病害状態を画像から分類するデータセットと解析モデルを開発・評価しており、植物病害フェノタイピング手法が中心である。

titleA comprehensive annotated image dataset for deep learning analysis of eggplant leaf diseases.
Reproduction assets foundThe paper's eggplant leaf disease image dataset (3116 annotated images, augmented to 10,000) is publicly deposited on Mendeley Data with a direct URL and DOI provided in the article.
Dataset · publicSeed Certification Agency, Ministry of Agriculture, Bangladesh, for his invaluable feedback and cooperation . Data source location Town/City/Region: Dhaka, Musnshigonj and Jhenaidah Sadar. Country: Bangladesh Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/5drkk544k8.1 Direct URL to data: https://data.mendeley.com/datasets/5drkk544k8/1Open asset ↗Mendeley Data · 10.17632/5drkk544k8.1lines:1-43
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published8 Oct 2025Data in briefCited by 3 · OpenAlex ↗

Cotton leaf image dataset for disease classification and health monitoring.

CottonField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Cotton, often referred to as "white gold" or the "king of fibers," is one of the most widely used natural fibers in the global textile industry, supporting approximately 250 million people worldwide. However, cotton plants suffer from a variety of diseases, particularly leaf diseases, which can significantly reduce the yield and fiber quality. To overcome this problem, we propose a carefully curated image dataset that enables research toward early and automated disease detection and health monitoring of cotton plants. The dataset comprises 1373 original and 4963 augmented high-resolution images of cotton leaves with healthy, damaged, and infected samples. The images were captured under different environmental conditions from plants grown at the Sher-e-Bangla Agricultural University in Dhaka, Bangladesh to provide natural variability and realism. The dataset considers four common cotton leaf diseases-Fusarium wilt, Alternaria leaf spot, Verticillium wilt, and bacterial blight-each labeled and classified to support machine learning applications. Captured from different angles and devices, the images have rich visual content that enables the development of strong deep learning models for disease classification. The dataset was designed to advance research relevant to precision agriculture by supporting early disease detection studies, crop health monitoring, and sustainable cotton-growing methods.

Why it matches plant phenotyping methods綿花葉の病害・健全状態を画像で分類するためのデータセットであり、植物の病害状態を直接観測する再利用可能なフェノタイピング資源が中心です。

abstractwe propose a carefully curated image dataset that enables research toward early and automated disease detection and health monitoring of cotton plants.
Reproduction assets foundThe paper is a data article describing a cotton leaf image dataset (1373 original + 4963 augmented images) for disease classification, publicly deposited on Mendeley Data with DOI 10.17632/t9hgvk2h9p.1 and a direct URL. This is the paper's own plant-phenotyping (leaf disease image) dataset and is directly actionable.
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/t9hgvk2h9p.1 Direct URL to data: https://data.mendeley.com/datasets/t9hgvk2h9p/1Open asset ↗Mendeley Data · 10.17632/t9hgvk2h9p.1lines:1-48
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 14 Sept 2026
Published6 Oct 2025Plant PhenomicsCited by 4 · OpenAlex ↗

3DPotatoTwin: a paired potato tuber dataset for 3D multi-sensory fusion

PotatoField / plotGrowth chamberLaboratory / benchtopPhotogrammetry / SfM / MVSLiDAR / point cloudRGB-D / ToFWhole plant / canopy / plot / fieldAnnotation / quality control2D/3D reconstruction

Accurate 3D phenotyping of agricultural produce remains challenging due to the trade-off between reconstruction quality and acquisition throughput in existing sensing technologies. While RGB-D cameras enable high-throughput scanning in operational settings like harvesting conveyors, they produce incomplete, low-quality 3D models. Conversely, close-range Structure-from-Motion (SfM) produces high-quality reconstructions but is not suitable for high-throughput field application. This study bridges this gap through 3DPotatoTwin , a paired dataset containing 339 tuber samples across three cultivars collected in Hokkaido, Japan. Our dataset uniquely combines: (1) conveyor-acquired RGB-D point clouds, (2) ground measurement, (3) SfM reconstructions under indoor controlled environment, and (4) aligned model pairs with transformation matrices. The multi-sensory alignment employs an semi-supervised pin-guided pipeline incorporating single-pin extraction and referencing, cross-strip matching, and binary-color-enhanced ICP, achieving 0.59 ​± ​0.11 ​mm registration accuracy. Beyond serving as a benchmark for 3D phenotyping algorithms, the dataset enables training of 3D completion networks to reconstruct high-quality 3D models from partial RGB-D point clouds. Meanwhile, the proposed semi-automated annotation pipeline has the potential to accelerate 3D dataset generation for similar studies. The presented methodology demonstrates broader applicability for multi-sensor data fusion across crop phenotyping applications. The dataset and pipeline source code are publicly available at HuggingFace and GitHub, respectively.

Why it matches plant phenotyping methodsジャガイモ塊茎の3D表現型計測を対象に、RGB-D・SfM・地上計測を統合したデータセット、位置合わせパイプライン、ベンチマークを開発しており、表現型取得手法が中心である。

abstractAccurate 3D phenotyping of agricultural produce remains challenging due to the trade-off between reconstruction quality and acquisition throughput in existing sensing technologies.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicAll the batch processing scripts mentioned in this section were provided in the 3dscan folder at Github (https://github.com/UTokyo-FieldPhenomics-Lab/PotatoScan/).Open asset ↗UTokyo-FieldPhenomics-Lab/PotatoScanhtml-lines:119-131
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Oct 2025Data in briefCited by 8 · OpenAlex ↗

PlantCity: A comprehensive image based on multi crop leaves in Pakistan.

AppleCherryCommon beanGrapevineMaizePearTomatoField / plotLeafClassification

The PlantCity dataset addresses significant agricultural yield losses in Pakistan from plant diseases. It provides 10,667 high-resolution images of leaves from 12 key crops: apple, apricot, bean, cherry, maize, fig, grape, loquat, pear, tomato, walnut, and persimmon. The images are organized into 52 classes (41 diseased and 11 healthy) and augmented to a total of 52,273 images. Data was collected in real-field conditions in Charsadda (34.15°N, 71.74°E, typical temperature 40-44 °C) and Chitral (35.85°N, 71.79°E, typical temperature 25-30 °C) from April to July 2023-2024. The dataset enables the development of deep learning models for automated disease classification and captures a range of environmental factors, including high temperatures that can exacerbate disease symptoms. It utilizes smartphone-based computer vision to facilitate early disease identification, thereby supporting precision farming and sustainable agriculture in Pakistan.

Why it matches plant phenotyping methods植物葉の病害状態を画像から分類するデータセットが研究の中心であり、植物病害フェノタイピング用の画像データセットとして適格です。

abstractThe PlantCity dataset addresses significant agricultural yield losses in Pakistan from plant diseases.
Reproduction assets foundThe paper is a Data in Brief article describing the PlantCity plant leaf image dataset (10,667 original images, 52 classes, 12 crops, collected in Pakistan). The dataset itself is the paper's core phenotyping asset and is publicly deposited on Mendeley Data with a direct URL provided in the article.
Dataset · publicon of diseases, pests, or environmental stress in plant leaves. Data source location Charsadda (Village Sarki) chosen for tomato and Chitral (Village Danin) for the other 11 crops, Khyber Pakhtunkhwa, Pakistan Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/w8kh2xkspx.2 Direct URL to data: https://data.mendeley.com/datasets/w8kh2xkspx/1 Related research article None 1 Value of the Data • The PlantCity dataset is comprehensive, consisting of 10,667 high-resolution images across 52 classes (41 diseased, 11 healthy) from 12 crop species, collected from Charsadda (tomato disease symptoms) and Danin Chitral (selected for its agro-climatic suitability for fruOpen asset ↗Mendeley Data · 10.17632/w8kh2xkspx.2lines:1-48
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published1 Oct 2025Data in BriefCited by 2 · OpenAlex ↗

A high-throughput phenotyping dataset for GWAS analysis of maize under combined drought and heat stress.

MaizeGrowth chamberWhole plant / canopy / plot / fieldMorphology / geometry measurementPhysiological trait estimationGrowth / time-series analysisGrowth / development / phenologyPhotosynthesis / fluorescenceStress response / tolerance

This dataset was generated to characterize the physiological and morphological mechanisms underlying tolerance and resilience to combined drought and heat stress using a panel of 106 Mediterranean maize inbred lines. To achieve this, high-throughput non-invasive phenotyping combined with genome-wide association analysis was applied to accurately capture the dynamic responses of the maize lines to stress and to dissect the genetic basis of maize tolerance and resilience. Two experiments were conducted under control (25/20 °C, 70 % field capacity (FC)) and stress conditions (35/25 °C, 30 % FC). Stress was applied from 18 to 32 DAS (days after sowing), followed by a recovery period under control conditions. Plants were grown under controlled air temperature and soil water content, and were harvested at 45 DAS. Throughout the cultivation period, multiple camera sensors captured images daily, allowing agronomic traits to be extracted for analysis. The dataset includes raw and processed images, phenotypic data obtained from these images, results of two photosynthesis related parameters, Genome-Wide Association Study (GWAS) results from one parameter as an example, and scripts used for data analysis. Additionally, metadata and a detailed description of the experimental setup are provided. This resource is suitable for researchers interested in stress phenotyping and quantitative genetics. It allows further exploration of genotype-by-environment interactions and integration with other omics datasets. The dataset provides a valuable foundation for studies aiming to understand and improve crop resilience to climate-related abiotic stresses.

Why it matches plant phenotyping methods植物の高スループット表現型取得を中心とするデータセットで、画像から農業形質を抽出するセンサー基盤、処理画像、表現型データ、解析スクリプトを提供しているため。

abstracthigh-throughput non-invasive phenotyping combined with genome-wide association analysis was applied to accurately capture the dynamic responses of the maize lines to stress
Reproduction assets foundThe authors deposited the paper's raw/processed phenotyping images, phenotypic and photosynthesis data, GWAS inputs/results, and R analysis scripts in the public e!DAL repository (DOI 10.5447/ipk/2025/8) in ISA-Tab/MIAPPE format.
Dataset · publicThe produced raw datasets and source code were uploaded to the e!DAL repository in ISA-Tab format (http://dx.doi.org/10.5447/ipk/2025/8) according to the MIAPPE standard.Open asset ↗e!DAL · 10.5447/ipk/2025/8html-lines:126-157
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published1 Oct 2025Data in BriefCited by 1 · OpenAlex ↗

Dataset of Ash gourd plant leaf images for detection and classification

Pumpkin / squashLeafClassificationObject detectionDisease symptoms / severity

The Ash Gourd dataset is valuable since it was collected from the diverse regions within the district of Dhaka in Bangladesh. This dataset represents one of the first attempts to document, elicit, and categorize the health conditions of Ash Gourd (Benincasa hispida) plants in Bangladesh based on healthy samples, aphid plurality, downy mildew, leaf curl, and leaf miner-infested categories. Ash Gourd is one of the region's most important vegetables because of its nutritional and economic value; thus, it is essential to know diseases' manifestation in the improvement of agricultural productivity. The Ash Gourd dataset contains 2676 images, structured into the five categories of Healthy, Aphid, Downy Mildew, Leaf Curl, and Leaf Miner. All images in all categories are raw which can be used flexibly according to the needs of analysis and model training. Concretely, the Healthy class consists of 803 images, while the four other classes contain 1,873 images. This structured way of collecting data will, in turn, enable deeper analysis and help construct machine learning models for disease classification, hence providing worthy insights into Ash Gourd plant health.

Why it matches plant phenotyping methodsアッシュゴード葉の画像データセットを構築し、植物の健康状態・病徴カテゴリを分類するための再利用可能なデータ資源を提供しており、植物病害状態の画像ベース表現型解析が中心です。

abstractThe Ash Gourd dataset contains 2676 images, structured into the five categories of Healthy, Aphid, Downy Mildew, Leaf Curl, and Leaf Miner.
Reproduction assets foundThe paper's own ash gourd leaf image dataset (2676 images, five classes) is publicly deposited on Mendeley Data with an explicit direct URL and DOI, matching the paper's phenotyping measurements.
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/zj4th6xvdp.2 Direct URL to data:https://data.mendeley.com/datasets/zj4th6xvdp/2Open asset ↗Mendeley Data · 10.17632/zj4th6xvdp.2html-lines:1-98
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published30 Sept 2025Cited by 0 · OpenAlex ↗

BudCAM: An Edge-Computing Camera System for Bud Detection in Muscadine Grapevines

GrapevineField / plotRGB / grayscaleObject detectionGrowth / development / phenology

Bud break is a critical phenological stage in muscadine grapevines, marking the start of the growing season and the increasing need for irrigation management. Real-time bud detection enables irrigation to match muscadine grape phenology, conserving water and enhancing performance. This study presents BudCAM, a low-cost, solar-powered, edge-computing camera system based on Raspberry Pi 5 and integrated with LoRa radio board, developed for real-time bud detection. Nine BudCAMs were deployed at Florida A&M University Center for Viticulture and Samll Fruit Research from mid February to mid March, 2024, monitoring three wine cultivars (A-27, noble, and Floriana) with three replicates each. Muscadine grape canopy images were captured every 20 minutes between 7:00 to 19:00, generating 2656 high-resolution (4656×3456 pixels) bud break images as database for bud detection algorithm development. The dataset was divided into 70% training, 15% validation, and 15% test. YOLOv11 models were trained using two primary strategies: a direct single-stage detector on tiled raw images and a refined two-stage pipeline that first identifies the grapevine cordon. Extensive evaluation of multiple model configurations identified top performers for both the single-stage (mAP@0.5=86.0%) and two-stage (mAP@0.5=85.0%) approaches. Further analysis revealed that preserving image scale via tiling was superior to alternative inference strategies like resizing or slicing. Field evaluations during the 2025 growing season confirmed the system’s effectiveness, with the two-stage model showing greater robustness to environmental noise like lens fog. A time-series filter smooths the raw daily counts to reveal a clear phenological trend for visualization. In its final deployment, the autonomous BudCAM system captures an image, runs inference on-device, and transmits the bud count in under three minutes, demonstrating a complete, field-ready solution for precision vineyard management.

Why it matches plant phenotyping methodsブドウの芽数・芽吹きという植物の表現型を、エッジカメラ、画像データセット、検出アルゴリズム、時系列処理で取得・推定するシステムを開発・評価しており、方法が研究の中心である。

abstractThis study presents BudCAM, a low-cost, solar-powered, edge-computing camera system based on Raspberry Pi 5 and integrated with LoRa radio board, developed for real-time bud detection.
Reproduction assets foundThe paper describes a public, custom-designed website dashboard that displays near-real-time bud detection counts from the BudCAM system (RG-trained Model 7) for the study's vines, including a specific sensor view (NP6, Mar–Apr 2025). This is a paper-specific public asset reproducing the paper's phenotyping outputs. No
Dataset · publicdetected images processed by CC-trained Model 8. Detected buds are highlighted with red bounding boxes, and the non-detected buds are highlighted with orange bounding boxes. Figure 11. Example dashboard view (NP6) showing raw counts at 30-minute intervals (Mar-Apr 2025). Results were produced by RG-trained Model 7. Available at https://phrec-irrigation.com/#/f/124/sensors/410.Preprints.org (www.preprints.org) | NOT PEER-REVIEWED | Posted: Posted: 30 September 2025doi:10.20944/preprints202509.2530.v1 © 2025 by the author(s). Distributed under a Creative Commons CC BY license.Open asset ↗pdf-raw-page:16 lines:1-9
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published30 Sept 2025Plant phenomics (Washington, D.C.)

Three-dimensional reconstruction of densely planted rice seedlings based on MultiView images.

RiceLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detection2D/3D reconstructionImage / point-cloud registrationGrowth / development / phenologyPlant / canopy height

Three-dimensional(3D) seedling reconstruction technology can provide critical technical support for monitoring plant growth, phenotyping high-throughput plants, and conducting precision agriculture. However, multiview image-based reconstruction methods, which rely on image registration and feature matching, are susceptible to issues such as similar textures and viewpoint differences, leading to matching errors and the loss of key structural information. This can result in local deficiencies and reduced accuracy in the reconstructed models. Therefore, to attain improved reconstruction accuracy under low-cost constraints, deep learning-based feature extraction and matching methods are employed in this study, the SuperPoint network is utilized to increase the robustness of the feature point detection and description processes, and the LightGlue algorithm is introduced to improve the accuracy and stability of matching. Additionally, to reduce the impact of shooting and platform jitter on image quality, a dedicated plant 3D reconstruction platform is designed and constructed, and a dataset of densely planted rice seedlings under light stress conditions is collected, comprising three factors (light quality, light quantity, and the photoperiod) ​× ​three levels, totaling nine groups. Experimental results show that the proposed method achieves optimal performance in terms of its point cloud completeness and reprojection error. The phenotypic parameters (e.g., plant height) extracted from the reconstruction data are strongly correlated with the actual measurements (R 2 ​= ​0.989, RMSE ​= ​4.54 ​mm), validating the potential of the proposed method for applications related to simulating plant growth processes, analyzing the effects of environmental factors (e.g., light), and optimizing crop cultivation schemes.

Why it matches plant phenotyping methodsマルチビュー画像によるイネ幼苗の3D再構成プラットフォームとデータセットを開発し、再構成精度および抽出形質を実測値と検証しており、表現型取得手法が研究の中心である。

abstractdeep learning-based feature extraction and matching methods are employed in this study
Reproduction assets foundThe paper's authors explicitly state their analysis code is publicly available on GitHub; the phenotype/image dataset is only available upon request, so it does not qualify as a public asset.
Code · publicThe code used in this study is available at https://github.com/Terrywewee/3D-reconstruction-of-densely-planted-rice-seedlings---superpoint-lightglue.git .Open asset ↗https://github.com/Terrywewee/3D-reconstruction-of-densely-planted-rice-seedlings---superpoint-lightglue.gitlines:433-485
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published30 Sept 2025Plant phenomics (Washington, D.C.)

SegPPD-FS: Segmenting plant pests and diseases in the wild using few-shot learning.

Field / plotSegmentationStress / disease detectionDisease symptoms / severity

Accurate segmentation of areas affected by pests and diseases is essential for precisely assessing the severity and spread of infections, thereby facilitating the development of effective management and intervention strategies. Obtaining high-quality pixel-level annotations for training deep learning models in agricultural environments poses considerable challenges. To overcome this limitation, the present work introduces a novel semantic segmentation approach (SegPPD-FS) that employs few-shot learning techniques to reduce annotation demands while effectively segmenting plant pests and diseases. The proposed SegPPD-FS comprises two key components: the similarity feature enhancement module (SFEM) and the hierarchical prior knowledge injection module (HPKIM). The SFEM refines foreground targets by employing a lightweight attention mechanism to mitigate irrelevant background interference in natural images and further enhances the discriminative capability of query features. The HPKIM is designed to address the difficulties associated with identifying pests and diseases that vary widely in terms of shape and size within field images, which is achieved through a hierarchical integration of multiscale contextual data into the query feature representations. In addition, this study constructed and publicly released a high-quality few-shot semantic segmentation (FSS) dataset that included 101 distinct categories of plant pests and diseases, which supports further research on the precise monitoring of plant health issues. The experimental results demonstrate that the proposed method achieves mIoU values of 71.19 ​% and 71.58 ​% with the 1-shot and 2-shot settings, respectively, on the released dataset. This performance surpasses that of other FSS techniques, such as SegGPT and PerSAM, providing a promising and label-efficient solution for pest and disease monitoring. The collected dataset, which focuses on plant pests and diseases, has been publicly released at https://doi.org/10.5281/zenodo.15114159, providing a valuable resource for evaluating various FSS techniques.

Why it matches plant phenotyping methods植物の病害・害虫による影響領域を画像からセグメンテーションし、被害の重症度・拡大を評価する手法を開発しており、植物状態の取得が中心です。公開データセットの構築・ベンチマークも含みます。

abstractthe present work introduces a novel semantic segmentation approach (SegPPD-FS) that employs few-shot learning techniques to reduce annotation demands while effectively segmenting plant pests and diseases
Reproduction assets foundThe paper publicly releases its SegPPD-101 pest/disease segmentation dataset (2263 pixel-annotated images, 101 categories) on Zenodo and its model weights via the authors' GitHub repository, both explicitly stated in the data availability statement.
Dataset · publicThe dataset used in this study is available at https://doi.org/10.5281/zenodo.15114159, and the model weights can be accessed at https://github.com/zihan303/SegPPD-FS.Open asset ↗Zenodo · 10.5281/zenodo.15114159html-lines:393-417
Code / dataset availability confirmedEurope PMC · bioRxiv · checked 6 Sept 2026
Published29 Sept 2025bioRxiv

Turning a new leaf: PhenoVision provides leaf phenology data at the global scale

RGB / grayscaleLeafAnnotation / quality controlClassificationGrowth / development / phenology

ABSTRACT Plant phenology dictates many aspects of community function and ecosystem dynamics. Yet, global phenology data are still limited, especially in areas lacking monitoring programs. Here we present a new data resource, PhenoVision–Leaf, which extends a computer-vision pipeline utilizing iNaturalist digital image vouchers to produce global-scale leaf phenophase data for deciduous, woody genera. We first discuss our implementation of a new human annotation framework for leaf phenology on iNaturalist, aligning with phenophase definitions used by the larger phenology community. We then showcase the use of 165,988 crowdsourced annotated records to train a Vision Transformer model with a two-stage regime to maximize accuracy across single- and multi-image records. This approach extends Phenovision from scoring individual images to aggregating at the iNaturalist record level, better aligning with human annotation processes. Post-hoc validation showed high performance for detecting present green and colored leaves (>98% accuracy), and reasonable accuracy for breaking leaf buds (>87% accuracy). Applying PhenoVision–Leaf to over 26 million iNaturalist records yielded 5.6 million record-level phenology observations across 6,500 species and 57 families, filling geographic and taxonomic gaps. These data, now accessible through the Phenobase portal, establish a foundation for near real-time monitoring of leaf phenology, supporting global-scale synthesis analyses.

Why it matches plant phenotyping methods葉のフェノロジー状態を画像から推定するコンピュータビジョン手法と、注釈・学習・検証・大規模データ生成基盤が研究の中心であるため。

abstractwe present a new data resource, PhenoVision–Leaf, which extends a computer-vision pipeline utilizing iNaturalist digital image vouchers to produce global-scale leaf phenophase data
Reproduction assets foundThe paper's PhenoVision–Leaf record-level leaf phenology dataset (5.6M machine-labeled observations) is publicly available via the Phenobase portal, and the underlying iNaturalist images used for training and machine labeling are available through the iNaturalist open data repository on AWS. No author analysis code or
Dataset · publicAll images associated with these records were downloaded using iNaturalist’s open data repository on AWS (https://registry.opendata.aws/inaturalist-open-data/).Open asset ↗pdf-page:4 lines:1-44
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published26 Sept 2025

Yield-Graph: Multi-stage Growth-aware Maize Yield Prediction via Graph Neural Networks

MaizeWhole plant / canopy / plot / fieldYield / biomass estimationGrowth / development / phenologyYield / yield components

Abstract Accurate yield prediction before maize harvest is crucial for advancing agricultural management and ensuring food security. Unlike conventional approaches that rely on phenotypes from a single growth stage, this study models multiple traits across different developmental stages, all targeting final yield, thereby uncovering their dynamic and cumulative contributions. We introduce Yield-Graph, an innovative framework that integrates multi-stage phenotypic data for yield prediction. The method employs a bipartite graph structure to impute missing trait values at each stage and leverages a hypergraph attention mechanism to capture high-order sample relationships. Comprehensive benchmark experiments demonstrate that Yield-Graph consistently outperforms traditional machine learning and graph-based models in both trait completion and yield prediction. Moreover, the framework exhibits strong robustness across growth stages, high adaptability to regional variations, and effective generalization across datasets. These findings highlight the potential of graph-enhanced multi-stage modeling for early-stage yield prediction, offering a scalable solution for precision agriculture and intelligent crop management.

Why it matches plant phenotyping methods多段階の植物形質を補完・統合し、収量という植物形質を予測するグラフ手法を開発・ベンチマークしており、形質取得・推定ワークフローが中心です。

abstractWe introduce Yield-Graph, an innovative framework that integrates multi-stage phenotypic data for yield prediction.
Reproduction assets foundThe paper's authors publicly release their Yield-Graph analysis code on GitHub; the phenotype datasets themselves are only available on request.
Code · publicthe manuscript. All authors read and approved the final manuscript. Data availability The datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request. Code availability The code developed to generate the results and analysis in this article is available at https://github.com/wjhhh2928/Yield-GraphOpen asset ↗https://github.com/wjhhh2928/Yield-Graphpdf-raw-page:14 lines:1-38
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published25 Sept 2025Data in briefCited by 0 · OpenAlex ↗

RoseLeafSet: Real-world leaf image dataset for AI-based agricultural solutions.

Field / plotLeafClassificationStress / disease detectionDisease symptoms / severity

This study highlights the growing significance of flowers, especially roses, in the global agricultural market, where they are cultivated for both personal enjoyment and commercial purposes. Among these, roses are considered one of the most popular and widely cultivated flowers. However, rose cultivators often encounter substantial challenges due to diseases that affect the plants, which can lead to significant economic losses in the agricultural sector. Timely and accurate detection of these diseases is crucial to mitigating their impact, potentially saving millions of dollars in crop losses. The dataset utilized in this research consists of 10,000 high-quality images collected from an initial set of 3113 images taken from several rose gardens located in Amin Model Town, Khagan, Ashulia, and Savar, Bangladesh. The data collection process spanned from October 30 to November 6, 2024. These images are categorized into four distinct classes: Healthy Leaf, Black Spot, Leaf Hole, and Dry Leaf, representing various stages of disease development in rose plants. The images were captured using a Vivo IQOO Z9x phone, ensuring high resolution and detailed imagery necessary for research analysis. This dataset serves as a valuable resource for researchers and developers working on creating efficient algorithms for the early and accurate identification of rose leaf diseases. By leveraging machine learning and image processing techniques, these algorithms could significantly enhance disease detection and prevention, helping to safeguard crops and reduce economic losses in the agricultural sector.

Why it matches plant phenotyping methodsバラ葉の病徴を画像データセットとして体系的に収集・分類し、植物病害状態の画像ベース推定を支援する研究であり、データセット構築が中心です。

titleRoseLeafSet: Real-world leaf image dataset for AI-based agricultural solutions.
Reproduction assets foundThe paper is a data descriptor for RoseLeafSet, a public rose leaf image dataset (3113 original images, augmented to 10,000) deposited on Mendeley Data with an explicit direct URL and DOI (10.17632/9g668bfhy5.3). This is a paper-specific, publicly available plant image dataset directly reproducing the paper's phenotypy
Dataset · publiclocation City: Amin Model Town, Khagan, Ashulia, Savar, Dhaka Country: Bangladesh. Local location: Shumi Nursery, Shetu Nursery, Bismillah Nursery etc. Geographical Location: 23 ° 53′ 2″ N and 90 ° 19′ 28″ E. Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/9g668bfhy5.3 Direct URL to data: https://data.mendeley.com/datasets/9g668bfhy5/3 Related research article None 1. Value of the Data • The dataset presented here, a collaborative effort of researchers and industry professionals, is suitable for training machine learning models for rose leaf disease classification and detection. This makes it a valuable resource for all of us, as we work together to devOpen asset ↗Mendeley Data · 10.17632/9g668bfhy5.3lines:1-48
Code / dataset availability confirmedEurope PMC · bioRxiv · checked 15 Sept 2026
Published19 Sept 2025bioRxiv

AngleCam V2: Predicting leaf inclination angles across taxa from daytime and nighttime photos

LiDAR / point cloudRGB / grayscaleLeafMorphology / geometry measurementObject detectionStress / disease detectionTrackingArchitecture / morphology / geometryLeaf traits

Understanding how plants capture light and maintain their energy balance is crucial for predicting how ecosystems respond to environmental changes. By monitoring leaf inclination angle distributions (LIADs), we can gain insights into plant behaviour that directly influences ecosystem functioning. LIADs affect radiative transfer processes and reflectance signals, which are essential components of satellite-based vegetation monitoring. Despite their importance, scalable methods for continuously observing these dynamics across different plant species throughout day-night cycles are limited. We present AngleCam V2, a deep learning model that estimates LIADs from both RGB and near-infrared (NIR) night-vision imagery. We compiled a dataset of over 4,500 images across 200 globally distributed species to facilitate generalization across taxa. Moreover, we developed a method to simulate pseudo-NIR imagery from RGB imagery to enable an efficient training of a deep learning model for tracking LIADs across day and night. The model is based on a vision transformer architecture with mixed-modality training using the RGB and the synthetic NIR images. AngleCam V2 achieved substantial improvements in generalization compared to AngleCam V1 (R 2 = 0.62 vs 0.12 on the same holdout dataset). Phylogenetic analysis across 100 genera revealed no systematic taxonomic bias in prediction errors. Testing against leaf angle dynamics obtained from multitemporal terrestrial laser scanning demonstrated the reliable tracking of diurnal leaf movements (R 2 = 0.61-0.75) and the successful detection of water limitation-induced changes over a 14-day monitoring period. This method enables continuous monitoring of leaf angle dynamics using conventional cameras, enabling applications in ecosystem monitoring networks, plant stress detection, interpreting satellite vegetation signals, and citizen science platforms for global-scale understanding of plant structural responses.

Why it matches plant phenotyping methods葉の傾斜角分布という植物形質を画像から推定する深層学習手法を開発し、大規模データセット、既存モデル比較、レーザースキャンによる検証、水ストレス下での追跡評価まで実施しており、フェノタイピング手法が研究の中心です。

abstractWe present AngleCam V2, a deep learning model that estimates LIADs from both RGB and near-infrared (NIR) night-vision imagery.
Reproduction assets foundThe paper's Data Availability Statement explicitly provides public access to the authors' analysis code (Anonymous GitHub), the phenotyping image/trait dataset (Zenodo), and the pretrained AngleCam V2 model weights (Zenodo). All three are paper-specific, public, and actionable.
Code · publicLK and TK conceived the ideas, designed the methodology, and led the analysis. TK, JP, RR, JF, LK, 26 and DL collected the data. LK and TK led the writing of the manuscript. All authors contributed 27 critically to the drafts and gave final approval for publication. 28 Data Availability Statement 29 The code is available here (https://anonymous.4open.science/r/AngleCamV2-2B38). The data 30 is available at (https://doi.org/10.5281/zenodo.17086253). The pretrained model is available 31 at (https://doi.org/10.5281/zenodo.17101166).32 Conflicts of Interest 33 All authors declare that they have no conflicts of interest. 34 2 . CC-BY 4.0 International license perpetuity. It is made available underOpen asset ↗anonymous.4open.science/r/AngleCamV2-2B38pdf-raw-page:2 lines:1-30
Dataset · publicK, JP, RR, JF, LK, 26 and DL collected the data. LK and TK led the writing of the manuscript. All authors contributed 27 critically to the drafts and gave final approval for publication. 28 Data Availability Statement 29 The code is available here (https://anonymous.4open.science/r/AngleCamV2-2B38). The data 30 is available at (https://doi.org/10.5281/zenodo.17086253). The pretrained model is available 31 at (https://doi.org/10.5281/zenodo.17101166).32 Conflicts of Interest 33 All authors declare that they have no conflicts of interest. 34 2 . CC-BY 4.0 International license perpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, whoOpen asset ↗zenodo · 10.5281/zenodo.17086253pdf-raw-page:2 lines:1-30
Model / weights · publicanuscript. All authors contributed 27 critically to the drafts and gave final approval for publication. 28 Data Availability Statement 29 The code is available here (https://anonymous.4open.science/r/AngleCamV2-2B38). The data 30 is available at (https://doi.org/10.5281/zenodo.17086253). The pretrained model is available 31 at (https://doi.org/10.5281/zenodo.17101166).32 Conflicts of Interest 33 All authors declare that they have no conflicts of interest. 34 2 . CC-BY 4.0 International license perpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for tOpen asset ↗zenodo · 10.5281/zenodo.17101166pdf-raw-page:2 lines:1-30
Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Published16 Sept 2025arXiv

WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory

MultimodalLiDAR / point cloudWhole plant / canopy / plot / fieldClassificationSegmentation

Street trees are vital to urban livability, providing ecological and social benefits. Establishing a detailed, accurate, and dynamically updated street tree inventory has become essential for optimizing these multifunctional assets within space-constrained urban environments. Given that traditional field surveys are time-consuming and labor-intensive, automated surveys utilizing Mobile Mapping Systems (MMS) offer a more efficient solution. However, existing MMS-acquired tree datasets are limited by small-scale scene, limited annotation, or single modality, restricting their utility for comprehensive analysis. To address these limitations, we introduce WHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset. Collected across two distinct cities, WHU-STree integrates synchronized point clouds and high-resolution images, encompassing 21,007 annotated tree instances across 50 species and 2 morphological parameters. Leveraging the unique characteristics, WHU-STree concurrently supports over 10 tasks related to street tree inventory. We benchmark representative baselines for two key tasks--tree species classification and individual tree segmentation. Extensive experiments and in-depth analysis demonstrate the significant potential of multi-modal data fusion and underscore cross-domain applicability as a critical prerequisite for practical algorithm deployment. In particular, we identify key challenges and outline potential future works for fully exploiting WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV/WHU-STree.

Why it matches plant phenotyping methods樹木の個体セグメンテーションと形態パラメータを含むマルチモーダルデータセットを構築し、ベンチマークする研究であり、植物個体の状態・形態抽出手法が中心である。

abstractWHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset.
Reproduction assets foundThe paper's core asset is the WHU-STree multi-modal street tree dataset (point clouds, panoramic images, 21,007 annotated tree instances, 50 species, height/DBH), which the authors state is publicly accessible via their GitHub organization WHU-USI3DV. The Zenodo DOIs in the reference list belong to cited prior datasets
Dataset · publicticular, we identify key challenges and outline potential future works for fully exploit- ing WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV /WHU-STree. Keywords: Deep learning, Tree inventory, Individual tree segmentation, Tree species classification, Multi-modal, Mobile mapping system 1. Introduction Street trees, vital to urban ecosystems, provide ecological benefits (e.g., shade (Kumar et al., 2024), air purification (Grundstrém and Pleijel, 2014), noise reductiOpen asset ↗WHU-STreepdf-raw-page:2 lines:1-35
Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Published15 Sept 2025arXiv

Cott-ADNet: Lightweight Real-Time Cotton Boll and Flower Detection Under Field Conditions

CottonField / plotFlowerFruitObject detection

Cotton is one of the most important natural fiber crops worldwide, yet harvesting remains limited by labor-intensive manual picking, low efficiency, and yield losses from missing the optimal harvest window. Accurate recognition of cotton bolls and their maturity is therefore essential for automation, yield estimation, and breeding research. We propose Cott-ADNet, a lightweight real-time detector tailored to cotton boll and flower recognition under complex field conditions. Building on YOLOv11n, Cott-ADNet enhances spatial representation and robustness through improved convolutional designs, while introducing two new modules: a NeLU-enhanced Global Attention Mechanism to better capture weak and low-contrast features, and a Dilated Receptive Field SPPF to expand receptive fields for more effective multi-scale context modeling at low computational cost. We curate a labeled dataset of 4,966 images, and release an external validation set of 1,216 field images to support future research. Experiments show that Cott-ADNet achieves 91.5% Precision, 89.8% Recall, 93.3% mAP50, 71.3% mAP, and 90.6% F1-Score with only 7.5 GFLOPs, maintaining stable performance under multi-scale and rotational variations. These results demonstrate Cott-ADNet as an accurate and efficient solution for in-field deployment, and thus provide a reliable basis for automated cotton harvesting and high-throughput phenotypic analysis. Code and dataset is available at https://github.com/SweefongWong/Cott-ADNet.

Why it matches plant phenotyping methods綿花の花・ボール認識を対象とする画像解析手法を開発し、データセット作成、外部検証、性能評価まで行っており、植物器官の表現型取得が中心である。

abstractWe propose Cott-ADNet, a lightweight real-time detector tailored to cotton boll and flower recognition under complex field conditions.
Reproduction assets foundThe paper explicitly states that its code and curated cotton boll/flower detection dataset (4,966 labeled images plus a 1,216-image external validation set) are publicly released at the authors' GitHub repository. The ultralytics repository is a generic third-party library, not a paper-specific asset.
Code · publicy 7.5 GFLOPs, maintaining stable performance under multi-scale and rotational variations. These results demonstrate Cott-ADNet as an accurate and efficient solution for in-field deployment, and thus provide a reliable basis for automated cotton harvesting and high-throughput phenotypic analysis. Code and dataset is available at https://github.com/SweefongWong/Cott-ADNet . † † footnotetext: ∗ * Corresponding author: cuij@wfu.edu Index Terms : cotton, cotton boll detection, lightweight object detection, rotational convolution 1 Introduction Cotton is one of the most critical economic crops worldwide, accounting for nearly 35% of global natural fiber production. It underpins industries such as Open asset ↗SweefongWong/Cott-ADNetlines:1-57
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published15 Sept 2025Data in briefCited by 10 · OpenAlex ↗

Money plant leaf (Epipremnum aureum): A comprehensive study of raw datasets with manual classification.

RGB / grayscaleLeafClassification

Money plants are widely recognized for their significant spiritual and air-purifying benefits. Research has proven that daily interaction with these vibrant indoor plants effectively reduces anxiety and stress. This paper introduces a robust dataset of 4302 healthy, unhealthy, combined, real and college premises images of money plants captured using smartphones. The dataset was collected from the educational hub, Dr. D. Y. Patil Institute of Technology, Pune campus, Maharashtra, India. Under controlled conditions, images were taken from a mobile device to ensure consistency and quality. From different angles and different backgrounds, images are captured. The aim of creating the dataset was to support researchers in achieving their objectives in the agricultural field and to explore our dataset so that it may be used for further research, investigation, and training of artificial intelligence models using our dataset.

Why it matches plant phenotyping methods植物の健康・不健康状態を画像で収集・分類したデータセットが論文の中心であり、病害・状態フェノタイピング用データセットとして適格です。

abstractThis paper introduces a robust dataset of 4302 healthy, unhealthy, combined, real and college premises images of money plants captured using smartphones.
Reproduction assets foundThis Data in Brief article describes its own public plant-phenotyping asset: a 4302-image money plant (Epipremnum aureum) leaf dataset with manual healthy/unhealthy classification, deposited on Mendeley Data (V4, DOI 10.17632/kd8hs7ch6t.4) with a Zenodo mirror and an authors' GitHub repository. The dataset is the paper
Dataset · publicRepository name : Epipremnum aureum (Money Plant Leaf) datasets Data identification number : Version: V4, 10.17632/kd8hs7ch6t.4 Direct URL to data: https://data.mendeley.com/datasets/kd8hs7ch6t/4Open asset ↗10.17632/kd8hs7ch6t.4lines:1-73
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 6 Sept 2026
Published5 Sept 2025Research SquareCited by 0 · OpenAlex ↗

SegFormer Inspired Multi Head Spectral Attention with Edge Gating light weight model for Leaf Area Segmentation

Field / plotMultispectral / hyperspectralLeafSegmentationLeaf traits

Abstract Accurate segmentation of leaf area is a critical task in plant phenotyping and precision agriculture, as it directly impacts yield estimation, disease monitoring, and weed management. Conventional Convolutional Neural Networks (CNNs), such as UNet and its variants, often struggle with capturing long range contextual dependencies and preserving fine structural boundaries, while pure transformer based architectures like the Vision Transformer (ViT) suffer from poor inductive bias and limited data efficiency. To overcome these challenges , we propose a SegFormer inspired model that integrates Edge Gated Multi Head Spectral Attention (EG MHSA) for robust leaf area segmentation. The spectral attention mechanism captures discriminative frequency domain representations across spectral bands, while the edge gating module enhances boundary preservation by adaptively fusing multiscale edge features. Evaluated on the benchmark CWFID dataset, the proposed model achieves superior performance with an F1score of 97.33%, IoU of 95.84%, and the lowest loss of 0.0395, outperforming UNet variants and transformer based baselines. Qualitative analysis further demonstrates its effectiveness in accurately delineating fine leaf boundaries under complex field conditions. The ablation results highlight the complementary contributions of spectral attention and edge gating in boosting segmentation performance. With its lightweight architecture, edge focused refinement, and strong generalization capability, the proposed approach sets a new benchmark for leaf area segmentation and provides a practical, scalable solution for agricultural applications.

Why it matches plant phenotyping methods葉面積の画像セグメンテーション手法を開発・ベンチマーク評価しており、植物フェノタイピングにおける形態形質抽出が中心である。

abstractAccurate segmentation of leaf area is a critical task in plant phenotyping and precision agriculture
Reproduction assets foundThe paper evaluates its leaf area segmentation model on the public CWFID dataset (60 field images with pixel-level annotations), and the authors explicitly state the datasets are publicly available at the cwfid GitHub repository. No author analysis code or trained model checkpoints are reported.
Dataset · publicThe datasets used in the study are publicly available in the repository: https://github.com/cwfid/Open asset ↗https://github.com/cwfid/pdf-page:22 lines:1-27
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published4 Sept 2025Data in briefCited by 0 · OpenAlex ↗

A labeled image dataset of common tomato diseases for classification and object detection.

TomatoGreenhouseFruitLeafStem / branchClassificationObject detectionDisease symptoms / severity

Computer vision has emerged as a critical enabler of sustainable production in protected agriculture by offering efficient and non-invasive crop disease diagnosis. The development of accurate disease recognition models relies heavily on the availability of high-quality image datasets. This study introduces a tomato disease image dataset collected in 2024 from greenhouse facilities within a modern agricultural park in Sichuan Province, China. The dataset comprises 1026 high-resolution images, including 417 images of viral disease, 82 images of gray mold, and 527 images of bacterial wilt, totaling approximately 2.78 GB. Captured under real-world greenhouse conditions and from multiple angles and distances, the images effectively capture multi-scale phenotypic disease features. Manual annotation was conducted using the LabelImg tool under the guidance of plant pathology experts, with labeled regions covering leaves, fruits, and stems. Annotation files are stored in XML format, each corresponding to a specific image. This dataset is well-suited for research in disease classification, object detection, and phenotyping, and supports deep learning model training and cross-crop transfer learning applications.

Why it matches plant phenotyping methodsトマト病害の症状を画像で捉え、分類・検出モデル用に専門家アノテーションした再利用可能なデータセットであり、植物病害状態の表現型取得が中心である。

abstractThe development of accurate disease recognition models relies heavily on the availability of high-quality image datasets.
Reproduction assets foundThe paper is a Data in Brief article describing a public tomato disease image dataset (1026 annotated images) deposited on Mendeley Data with a direct URL and DOI, matching an allowed URL exactly.
Dataset · publicwas conducted at the Modern Agricultural Science and Technology Innovation Demonstration Park of the Sichuan Academy of Agricultural Sciences (30.7797° N, 104.2082° E), located in Sichuan Province, China. Data accessibility Repository name: Mendeley Data Data identification number: DOI: 10.17632/c2×8rynybg.1 Direct URL to data: https://data.mendeley.com/datasets/c2×8rynybg/1 Related research article None. 1 Value of the Data The dataset contains 1026 annotated images of tomato plants exhibiting three major disease types, collected in 2024 from greenhouse environments in Sichuan’s Modern Agricultural Demonstration Park. Plant pathology specialists manually labeled all samples. Its technical sOpen asset ↗Mendeley Data · 10.17632/c2×8rynybg.1lines:1-52
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 6 Sept 2026
Published2 Sept 2025aBIOTECHCited by 1 · OpenAlex ↗

FHBDSR-Net: automated measurement of diseased spikelet rate of Fusarium Head Blight on wheat spikes.

WheatRGB / grayscalePanicle / ear / spikeObject detectionDisease symptoms / severity

) disease that threatens global food security, requires precise quantification of diseased spikelet rate (DSR) as a phenotypic indicator for resistance breeding. Most techniques for measuring DSR rely on manual spikelet-by-spikelet observation and counting, which is inefficient and destructive. Although deep learning offers great promise for automated DSR measurement, existing intelligent detection algorithms are hampered by the lack of spikelet-level annotated data, insufficient feature representation for diseased spikelets, and weak spatial encoding of densely arranged spikelets. To address these challenges, we constructed a dataset of 620 high-resolution RGB images of wheat spikes with 5,222 spikelet-level annotations to systematically analyze spikelet size distributions to fill small-object detection data gaps in this field. We designed FHBDSR-Net, a light framework for automated DSR measurement centered on diseased spikelet detection, which features (1) multi-scale feature enhancement architecture that dynamically combines lesion textures, morphological features, and lesion-awn contrast through adaptive multi-scale kernels to suppress background noise; (2) the Inner-EfficiCIoU loss function to reduce small-target localization errors in dense contexts; and (3) a scale-aware attention module using dilated convolutions and self-attention to encode multi-scale pathological patterns and spatial distributions to enhance dense spikelet resolution. FHBDSR-Net detected diseased spikelets with an average precision of 93.8% with a lightweight design of 7.2 M parameters. The results were strongly correlated with expert evaluations, with a Pearson correlation coefficient of 0.901. Our method is suitable for deployment on resource-constrained mobile devices, facilitating portable plant phenotyping and smart breeding.

Why it matches plant phenotyping methodsコムギ穂の罹病小穂率という植物病害形質を画像から自動推定する手法を開発し、データセット構築と専門家評価による検証を行っており、フェノタイピング手法が中心である。

abstractrequires precise quantification of diseased spikelet rate (DSR) as a phenotypic indicator for resistance breeding.
Reproduction assets foundThe paper's Data availability statement explicitly deposits both the spikelet-level annotated wheat spike image dataset (620 RGB images, 5,222 annotations) and the FHBDSR-Net analysis code in a public GitHub repository under the authors' account, matching an allowed URL.
Dataset · publicThe dataset and code generated in this study are available at https://github.com/WeizhenLiuBioinform/Wheat-FHB-DSR-Measurement .Open asset ↗WeizhenLiuBioinform/Wheat-FHB-DSR-Measurementlines:901-961
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published23 Aug 2025Data in briefCited by 1 · OpenAlex ↗

TTADDA-UAV: A multi-season RGB and multispectral UAV dataset of potato fields collected in Japan and the Netherlands.

PotatoAerial / UAVField / plotRGB / grayscaleMultispectral / hyperspectralWhole plant / canopy / plot / fieldYield / yield components

The Transition to a Data-Driven Agriculture (TTADDA) project focuses on advancing the shift toward high-tech, circular agriculture. By developing cutting-edge sensor technologies and AI-driven tools, the project aims to boost productivity through a data-centric potato production system that supports circular agricultural practices. Potato phenotyping is crucial for creating high-yielding, resilient, and sustainable potato crops, which are essential in global food systems. Specifically in the Netherlands, the global leader in seed potato production and Japan that produces certified seed potatoes under strict quality controls and phytosanitary regulations. A multi-season drone dataset from five potato trials-three in Japan and two in the Netherlands was collected. Each trial field was divided into small plots, each planted with a specific cultivar to assess varietal performance. Data included drone imagery (RGB and multispectral), manual yield and ground coverage measurements, and weather data. The combination of sensor versatility, diverse potato varieties, and varying climate and soil conditions between Japan and the Netherlands makes this dataset highly valuable and potentially reusable for a wide range of applications. Using MIAPPE for this dataset ensures consistent, clear documentation of sensors, varieties, and conditions, making the data findable, reusable, and easy to integrate with other studies. It also supports reproducibility and automated analysis across the multi-location trials.

Why it matches plant phenotyping methodsジャガイモの表現型解析を目的としたUAV RGB・マルチスペクトル画像と圃場測定を含む、多季節・多地点の再利用可能なデータセットの構築・標準化が中心である。

abstractA multi-season drone dataset from five potato trials-three in Japan and two in the Netherlands was collected.
Reproduction assets foundThe paper is a data descriptor for the TTADDA-UAV potato phenotyping dataset (RGB/multispectral orthomosaics, DSMs, yield, ground coverage, weather) publicly deposited on 4TU.ResearchData with a DOI, plus an authors' GitHub repository for loading the MIAPPE-formatted data.
Dataset · publicThe dataset is part of the following collection: Data identification number: doi.org.10.4121/936b5772–09fc–4856–983d-1f9cc2f38d15 Direct URL to data: ( https://doi.org/10.4121/936b5772-09fc-4856-983d-1f9cc2f38d15 ) The collection consist of metadata, and five related studies: TTADDA_NARO_2021, TTADDA_NARO_2022, TTADDA_NARO_2023, TTADDA_WUR_2022, TTADDA_WUR_2023 To visualise the metadata and download the dataset we recommend the following GIT: https://github.com/NPEC-NL/MIAPPE_TTADDA_dataset Related research article None 1 Value of the DOpen asset ↗10.4121/936b5772-09fc-4856-983d-1f9cc2f38d15lines:57-83
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 Aug 2025Data in briefCited by 2 · OpenAlex ↗

A comprehensive dataset of rice leaf images for disease detection using machine learning.

RiceLeafStress / disease detectionDisease symptoms / severity

This manuscript presents a comprehensive, expert-annotated dataset comprising 19,000 rice leaf images, including 2,753 original images and 16,247 augmented images, sourced from the Bangladesh Rice Research Institute (BRRI). The dataset includes seven disease classes: Healthy (603 original images), Rice Blast (696 original images), Scald (421 original images), Leaf-folder Injury (247 original images), Insect Infestation (281 original images), Rice Stripes (266 original images), and Tungro Disease (239 original images). These images, captured under varying environmental conditions using smartphone cameras, accurately reflect real-world conditions. The images have been meticulously annotated by agronomy experts for reliable disease labeling. To enhance dataset diversity, data augmentation methods such as rotation, scaling, brightness adjustment, and horizontal flipping were systematically applied, expanding the dataset by creating additional variants from the original images. The dataset serves as a rich resource for developing machine learning models for the automatic detection of rice diseases. This initiative aims to enable early disease detection, promote sustainable farming practices, and improve food security, particularly in rice-dependent developing countries.

Why it matches plant phenotyping methodsイネ葉画像を用いて病害状態を表現型として扱う、専門家注釈付きデータセットの構築・提供が中心であり、植物病害フェノタイピング手法の基盤となる。

abstractThis manuscript presents a comprehensive, expert-annotated dataset comprising 19,000 rice leaf images
Reproduction assets foundThe paper is a Data in Brief article describing a rice leaf disease image dataset (19,000 images, 7 classes) publicly deposited on Mendeley Data with DOI 10.17632/vwv3nry3wr.1. This is the paper's own phenotyping image dataset and is directly actionable. No separate analysis code or trained model checkpoint is reported
Dataset · publicData accessibility Repository name: Mendeley Data Data identification number: 10.17632/vwv3nry3wr.1 Direct URL to data: https://data.mendeley.com/datasets/vwv3nry3wr/1Open asset ↗Mendeley Data · 10.17632/vwv3nry3wr.1lines:1-60
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published12 Aug 2025Data in briefCited by 1 · OpenAlex ↗

RoseLeafInsight: A high-resolution image dataset for rose leaf disease recognition.

Field / plotLeafClassificationStress / disease detectionDisease symptoms / severity

The Rose (genus Rosa) has become a significant factor in the Bangladeshi flower industry, both in terms of exports and local consumption. However, rose farming in this country faces serious challenges due to diseases affecting its leaves, which weaken the plants and result in lower flower yields and financial losses for farmers. Rosa (genus Rosa) is one of the most attractive and commercially valuable flower genera. However, agricultural rose production faces several challenges, such as pesticide resistance, which affects plant growth and results in a reduced quantity and quality of healthy flowers. Several natural factors also cause interference with rose production. Most farmers involved in this industry have limited education, which hinders their ability to identify early-stage rose-leaf disease solely through visual inspection. Furthermore, limited communication with agricultural experts exacerbates the situation, leading to delayed interventions and economic losses. This study presents the rose leaf disease dataset, which would help enhance disease tracking, diagnosis, and research in roses. From October 2024 to January 2025, large-scale field surveys were conducted to capture quality images for each condition class in rose leaves. In this paper, four classes comprise 'Black Spot,' 'Insect Hole,' 'Yellow Mosaic Virus,' and 'Healthy,' representing different stages in disease progression. There are 3,228 original images, categorized as follows: Black Spot (409), Insect Hole (453), Yellow Mosaic Virus (680), and Healthy (1,686). During the pre-processing stage, the images are resized to 3000×3000 pixels, and low-quality, duplicate, or irrelevant images are removed to ensure high quality. We have employed various augmentation techniques, including rotation, flipping, contrast adjustment, blurring, shearing, zooming, and noise addition, to increase the dataset size and enhance model generalization. Datasets like this one are in high demand for agricultural research, leading to improved disease management and increased yields. These goals can be achieved through high-accuracy machine-learning models for early disease detection and cause identification. This gives the farmers more time to take necessary actions for disease prevention and pest control. This tech-based system combines the field of agriculture with the cutting edge of computer science and AI, making precision agriculture even more effective and efficient. Our dataset is designed to meet the need for data to train these models and provide a baseline benchmark for disease detection in our specific crop, the Rose. Improvements in different generations of models, as well as numerous other forms of scientific advancements, can lead to further increases in efficiency and ultimately result in better, smarter farms. In our initial testing for categorizing rose leaves, we employed two well-known transfer learning models. Among them, MobileNetV2 performed exceptionally well, achieving an accuracy of 96.79% in image classification. This dataset can be integrated with innovative farming equipment, such as drones and sensors, to monitor large fields in real-time. This dataset serves as a benchmark for training deep learning models, enabling enhanced automated monitoring and decision-making in precision agriculture.

Why it matches plant phenotyping methodsバラ葉の病徴を画像で分類する大規模データセットとベンチマークを構築しており、植物の病害状態を直接推定する画像ベース手法が中心である。

titleRoseLeafInsight: A high-resolution image dataset for rose leaf disease recognition.
Reproduction assets foundThe paper's own rose leaf disease image dataset (3,228 original images plus processed/augmented versions) is publicly deposited on Mendeley Data with an explicit direct URL and DOI, matching an allowed URL.
Dataset · publicRepository name: Mendeley Data Data identification number: 10.17632/8chrjdxn79.1 Direct URL to data: https://data.mendeley.com/datasets/8chrjdxn79/2 The dataset is publicly available and can be accessed via the provided Mendeley Data repository link.Open asset ↗Mendeley Data · 10.17632/8chrjdxn79.1lines:31-66
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published12 Aug 2025Frontiers in plant scienceCited by 3 · OpenAlex ↗

LCAMNet: a lightweight model for apple leaf disease classification in natural environments.

AppleField / plotLeafClassificationDisease symptoms / severity

Apple leaf diseases severely affect the quality and yield of apples, and accurate classification is crucial for reducing losses. However, in natural environments, the similarity between backgrounds and lesion areas makes it difficult for existing models to balance lightweight design and high accuracy, limiting their practical applications. In order to resolve the aforementioned problem, this paper introduces a lightweight converged attention multi-branch network named LCAMNet. The network integrates depthwise separable convolutions and structural re-parameterization techniques to achieve efficient modeling. To avoid feature loss caused by single downsampling operations, a dual-branch downsampling module is designed. A multi-scale structure is introduced to enhance lesion feature diversity representation. An improved triplet attention mechanism is utilized to better capture deep lesion features. Furthermore, a dataset named SCEBD is constructed, containing multiple common disease types and interference factors under natural environments, realistically reflecting orchard conditions. Experimental results show that LCAMNet achieves 92.60% accuracy on the SCEBD and 95.31% on a public dataset, with only 0.03 GFLOPs and 1.30M parameters. The model maintains high accuracy while remaining lightweight, enabling effective apple leaf disease classification in natural environments on devices with limited resources.

Why it matches plant phenotyping methodsリンゴ葉の病徴を画像から分類する軽量モデルを開発し、自然環境データセットを構築・評価しており、植物病害状態の画像ベース表現型推定が中心である。

abstractthis paper introduces a lightweight converged attention multi-branch network named LCAMNet.
Reproduction assets foundThe paper's data availability statement links three public image datasets directly used in its experiments: the FGVC8 Plant Pathology 2021 Kaggle dataset, the AppleLeaf9 GitHub dataset, and the ATLDSD dataset on ScienceDB. No author analysis code or trained model is released, and the self-constructed SCEBD has no own公开
Dataset · publicce Foundation Project (No. 2024MS06002), the Inner Mongolia Autonomous Region universities innovative research team project (No. NMGIRT2313) and the Inner Mongolia Natural Science Foundation Project (No. 2025ZD012). Data availability statement Publicly available datasets were analyzed in this study. This data can be found here: https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8 ; https://github.com/JasonYangCode/AppleLeaf9 ; https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578 . Author contributions YJ: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. HL: Funding acquisition, Resources,Open asset ↗plant-pathology-2021-fgvc8 · plant-pathology-2021-fgvc8lines:727-753
Dataset · publicomous Region universities innovative research team project (No. NMGIRT2313) and the Inner Mongolia Natural Science Foundation Project (No. 2025ZD012). Data availability statement Publicly available datasets were analyzed in this study. This data can be found here: https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8 ; https://github.com/JasonYangCode/AppleLeaf9 ; https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578 . Author contributions YJ: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. HL: Funding acquisition, Resources, Writing – review & editing. XF: Project administration, SupervisOpen asset ↗JasonYangCode/AppleLeaf9 · JasonYangCode/AppleLeaf9lines:727-753
Dataset · publicteam project (No. NMGIRT2313) and the Inner Mongolia Natural Science Foundation Project (No. 2025ZD012). Data availability statement Publicly available datasets were analyzed in this study. This data can be found here: https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8 ; https://github.com/JasonYangCode/AppleLeaf9 ; https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578 . Author contributions YJ: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. HL: Funding acquisition, Resources, Writing – review & editing. XF: Project administration, Supervision, Writing – review & editing. BW: Project aOpen asset ↗0e1f57004db842f99668d82183afd578lines:727-753
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published10 Aug 2025Data in briefCited by 0 · OpenAlex ↗

RGB image dataset for okra maturity classification to enhance agricultural quality and market readiness.

Laboratory / benchtopRGB / grayscaleFruitClassificationGrowth / development / phenology

Okra is a highly nutritious farming product that combats malnutrition issues while supporting sustainable agricultural methods. In order to keep its quality and versatility in preparation applications, it is important to classify its maturity stages into under-mature, mature, and over-mature categories. Classification is vital to identify the best time to harvest, satisfy market demands, minimize post-harvest losses, and optimize cooking uses. This data in brief uses a non-destructive approach to okra maturity classification based on a dataset of images taken under controlled illumination using a standard RGB camera. The dataset contains okra samples that were sourced from various farms and vegetable markets, ensuring that it encapsulates the natural variability found in real-world farm and market environments. The availability of such a large dataset enables the creation of precise classification models that can assist farmers in optimizing the time of harvest, fulfilling consumers' requirements, and improving market results. The research has great relevance to promoting agricultural quality evaluation and boosting market readiness using non-invasive techniques.

Why it matches plant phenotyping methodsRGB画像データセットによるオクラ果実の成熟段階分類が研究の中心であり、植物器官の状態を画像から推定する再利用可能なフェノタイピング資源に該当する。

titleRGB image dataset for okra maturity classification to enhance agricultural quality and market readiness.
Reproduction assets foundThe paper is a Data in Brief article whose core contribution is a public RGB okra image dataset (364 images across three maturity classes) deposited on Mendeley Data, directly serving as the paper's phenotyping image asset. No separate analysis code repository is described.
Dataset · publicVellore Institute of Technology - Chennai Campus. City/Country: Chennai, India. Latitude and longitude for collected samples/data: (12.8406° N, 80.1534° E), Vellore Institute of Technology - Chennai. Data accessibility Repository name: Okra Image Dataset Data identification number: DOI: 10.17632/jmhz4826f2.1 Direct URL to data: https://data.mendeley.com/datasets/jmhz4826f2/1 Related research article [ 1 ] 1 Value of the Data • Agricultural quality assessment is advancing through non-invasive methods that include the use of RGB image analysis for effective determination of okra maturity stages. • Utilization of ML and DL methods in recognizing visual characteristics, i.e., color, texture, andOpen asset ↗10.17632/jmhz4826f2.1lines:1-56
Code / dataset availability confirmedOpenAlex · checked 14 Sept 2026
Published8 Aug 2025Journal of Telecommunications and Information TechnologyCited by 2 · OpenAlex ↗

Enhancing Leaf Area Segmentation by Using Attention Gates and Knowledge Distillation in UNet Architecture

SunflowerLeafSegmentationLeaf traits

Accurate segmentation of leaf regions plays a vital role in plant phenotyping and agricultural analysis. This paper presents AKDUNet, a lightweight UNet-based architecture that integrates attention gates and knowledge distillation to improve segmentation performance while minimizing computational complexity. The architecture replaces traditional skip connections with attention gates to focus on salient spatial features and employs a two-stage training pipeline, where a compact student model learns from a deeper teacher model using a tailored distillation loss function. AKDUNet is evaluated on two benchmark datasets (CWFID and Sunflower) and outperforms a range of state-of-the-art models, including UNet++, Inception UNet, VGG-based UNets, SDUNet, INSCA UNet, and SegFormer. Ablation studies confirm the advantages of attention modules, and qualitative analyses using Grad-CAM visualizations reveal the model's ability to effectively focus on crucial leaf structures. The results demonstrate that AKDUNet is not only computationally efficient but also highly accurate, making it suitable for real-time deployment in resource-constrained agricultural environments.

Why it matches plant phenotyping methods植物の葉領域を抽出する画像セグメンテーション手法の開発とベンチマーク評価が中心であり、葉面積などの表現型取得に直接利用できる。

abstractAccurate segmentation of leaf regions plays a vital role in plant phenotyping and agricultural analysis.
Reproduction assets foundThe paper evaluates AKDUNet leaf segmentation on the public CWFID dataset (and a Sunflower dataset), and the acknowledgments explicitly state the datasets are publicly available at the authors' cited repository URL. No author analysis code or trained model checkpoints are released.
Code / dataset availability confirmedEurope PMC · bioRxiv · checked 14 Sept 2026
Published7 Aug 2025bioRxivCited by 1 · OpenAlex ↗

The Tonoplast Topology Index - a new metric for describing vacuole organization

ArabidopsisLaboratory / benchtopMicroscopyRootMorphology / geometry measurementArchitecture / morphology / geometry

Background The plant vacuole arises by orchestrated interplay of membrane trafficking, cytoskeletal rearrangements and a variety of signalling pathways. In the root, the characteristic large central vacuole develops by endomembrane reorganization occurring mainly in the transition zone. The vacuole’s bounding membrane - the tonoplast - can be visualized in vivo using fluorescent protein markers, allowing for quantitative analysis of confocal microscopy images. Tonoplast organization can thus serve as a sensitive indicator of changes to any of the processes involved in vacuole biogenesis. The Vacuolar Morphology Index (VMI) is widely accepted as a quantitative measure of vacuole structure. However, this metric has two drawbacks - it only reflects the size of the largest vacuolar compartment (missing therefore possible differences in the organization of smaller compartments), and its determination is labor intensive, limiting its use on large datasets. Results We developed an alternative metric for describing vacuole organization, named the Tonoplast Topology Index (TTI), which overcomes the above-mentioned shortcomings of the VMI. We compared the performance of our protocol with VMI on a simulated dataset and on real data. To validate the methods’ performance, we used it to confirm the previously reported differences in vacuole shape and size between Arabidopsis thaliana roots grown on the surface of an agar medium compared to those embedded inside the agar. Both VMI and TTI could efficiently detect the relatively subtle changes in vacuole organization depending on the position of the root in the agar, and provided correlated results. However, only TTI produced data with close to normal value distribution, simplifying subsequent statistical evaluation. Conclusions We present the protocol for TTI determination as a two-stage semi-automated procedure involving microscopic image analysis employing an ImageJ macro and subsequent processing of numeric data in the Jupyter Notebook environment, together with benchmarking image data. Since this implementation is freeware-based, platform-independent and (relatively) user-friendly, we hope it will find its use as a high throughput, added value alternative to the VMI metric.

Why it matches plant phenotyping methods植物の液胞構造を定量化する新規指標と半自動画像解析プロトコルを開発し、既存指標との比較・実データおよびシミュレーションによる検証、ベンチマークデータを提示しており、表現型取得・抽出法が中心である。

abstractWe developed an alternative metric for describing vacuole organization, named the Tonoplast Topology Index (TTI), which overcomes the above-mentioned shortcomings of the VMI.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicthe software tool generated here are also available at https://github.com/GeorgeCaldarescu/TTI-Open asset ↗GeorgeCaldarescu/TTI-pdf-page:9 lines:1-52
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published6 Aug 2025Plant phenomics (Washington, D.C.)Cited by 12 · OpenAlex ↗

The Global Wheat Full Semantic Organ Segmentation (GWFSS) dataset.

WheatField / plotPanicle / ear / spikeLeafSegmentation

Computer vision is increasingly used in farmers' fields and agricultural experiments to quantify important traits. Imaging setups with a sub-millimeter ground sampling distance enable the detection and tracking of plant features, including size, shape, and colour. Although today's AI-driven foundation models segment almost any object in an image, they still fail for complex plant canopies. To improve model performance, the global wheat dataset consortium assembled a diverse set of images from experiments around the globe. After the head detection dataset (GWHD), the new dataset targets a full semantic segmentation (GWFSS) of organs (leaves, stems and spikes) covering all developmental stages. Images were collected by 11 institutions using a wide range of imaging setups. Two datasets are provided: i) a set of 1096 diverse images in which all organs were labelled at the pixel level, and (ii) a dataset of 52,078 images without annotations available for additional training. The labelled set was used to train segmentation models based on DeepLabV3Plus and Segformer. Our Segformer model performed slightly better than DeepLabV3Plus with a mIOU for leaves and spikes of ca. 90 ​%. However, the precision for stems with 54 ​% was rather lower. The major advantages over published models are: i) the exclusion of weeds from the wheat canopy, ii) the detection of all wheat features including necrotic and senescent tissues and its separation from crop residues. This facilitates further development in classifying healthy vs. unhealthy tissue to address the increasing need for accurate quantification of senescence and diseases in wheat canopies.

Why it matches plant phenotyping methods小麦器官の画素レベルセグメンテーション用データセットを構築し、モデル性能を検証する研究であり、植物形質抽出のための画像解析手法が中心です。

abstractThe labelled set was used to train segmentation models based on DeepLabV3Plus and Segformer.
Reproduction assets foundThe paper's GWFSS wheat organ segmentation dataset (1096 pixel-labelled images plus 52,078 unlabelled images, subset/imaging-setup metadata) and the benchmark segmentation model are publicly deposited in the ETH Research Collection and mirrored on Hugging Face, with links also listed on the Global Wheat site.
Dataset · publicThe full dataset (GWFSS_v1.0_full) including the 1096 ground-truth labelled images (GWFSS_v1.0_labelled), the descriptions of the datasets (GWFSS_v1.0_subsets.csv) and imaging setups (GWFSS_v1.0_imaging_setups.csv) is available in the ETH research collection (https://doi.org/10.3929/ethz-b-000734546)Open asset ↗ETH research collection · 10.3929/ethz-b-000734546html-lines:1006-1041
Dataset · publicTo facilitate access, the labelled data and the benchmark model will also be available at (https://huggingface.co/datasets/GlobalWheat/GWFSS_v1.0).Open asset ↗huggingface · GlobalWheat/GWFSS_v1.0html-lines:1129-1192
Dataset · publicLinks to these datasets can be found at: https://www.global-wheat.com/gwfss.html.Open asset ↗html-lines:1006-1041
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published5 Aug 2025Frontiers in plant scienceCited by 3 · OpenAlex ↗

BiSeNeXt: a yam leaf and disease segmentation method based on an improved BiSeNetV2 in complex scenes.

YamLeafSegmentationDisease symptoms / severity

Introduction Yam is an important medicinal and edible crop, but its quality and yield are greatly affected by leaf diseases. Currently, research on yam leaf disease segmentation remains unexplored. Challenges like leaf overlapping, uneven lighting and irregular disease spots in complex environments limit segmentation accuracy. Methods To address these challenges, this paper introduces the first yam leaf disease segmentation dataset and proposes BiSeNeXt, an enhanced method based on BiSeNetV2. Firstly, dynamic feature extraction block (DFEB) enhances the precision of leaf and disease edge pixels and reduces lesion omission through dynamic receptive-field convolution (DRFConv) and pixel shuffle (PixelShuffle) downsampling. Secondly, efficient asymmetric multi-scale attention (EAMA) effectively alleviates the problem of lesion adhesion by combining asymmetric convolution with a multi-scale parallel structure. Finally, PointRefine decoder adaptively selects uncertain points in the image predictions and refines them point-by-point, producing accurate segmentation of leaves and spots. Results Experimental results indicated that the approach achieved a 97.04% intersection over union (IoU) for leaf segmentation and an 84.75% IoU for disease segmentation. Compared to DeepLabV3+, the proposed method improves the IoU of leaf and disease segmentation by 2.22% and 5.58%, respectively. Additionally, the FLOPs and total number of parameters of the proposed method require only 11.81% and 7.81% of DeepLabV3+, respectively. Discussion Therefore, the proposed method can efficiently and accurately extract yam leaf spots in complex scenes, providing a solid foundation for analyzing yam leaves and diseases.

Why it matches plant phenotyping methodsヤム葉と病斑を画像から分割・抽出する手法とデータセットを開発し、性能比較まで行っており、植物の病害状態を取得する方法が中心である。

abstractthis paper introduces the first yam leaf disease segmentation dataset and proposes BiSeNeXt, an enhanced method based on BiSeNetV2.
Reproduction assets foundThe paper's authors publicly released their self-constructed yam leaf disease segmentation dataset (1,097 annotated images of anthracnose, brown spot, and gray spot) via a Google Drive link in the Data availability statement. No code or trained model deposit is explicitly stated.
Dataset · publicThe datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: https://drive.google.com/drive/folders/1_ojcb_84TMbkZwYfm0dgsL1NjiGw7GRF?usp=sharing .Open asset ↗lines:1046-1093
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published5 Aug 2025Data in briefCited by 7 · OpenAlex ↗

AI-MedLeafX: a large-scale computer vision dataset for medicinal plant diagnosis.

LeafClassificationStress / disease detectionDisease symptoms / severity

This study presents a large, meticulously curated and manually validated dataset aimed at classifying leaf quality into five critical categories: Healthy, Bacterial Spot, Shot Hole, Yellow, and Powdery Mildew. The dataset encompasses four distinct plant species-Cinnamomum Camphora (Camphor), Terminalia Chebula (Haritaki), Moringa Oleifera (Sojina), and Azadirachta Indica (Neem)-each represented across three or four disease categories, depending on observed symptoms and final number of classes is thirteen (13 classes). Data collection was conducted between November 1, 2024, and January 5, 2025, utilizing four different mobile cameras to ensure diversity in image resolution, lighting, and environmental conditions. The original dataset comprised 10,858 high-resolution images, which were subsequently expanded to 65,148 through the application of six comprehensive data augmentation techniques, including rotations (45°, 60°, and 90°), horizontal flipping, zooming and brightness adjustment. All images were standardized to 512×512 pixels to ensure uniformity and seamless compatibility with machine learning and computer vision models. This enriched dataset serves as a crucial resource for the development of automated plant disease detection systems and supports advancements in precision agriculture. It not only addresses the pressing need for scalable, high-quality data in agricultural research but also establishes a solid foundation for benchmarking novel deep learning architectures. By enabling more accurate and efficient leaf disease classification, the dataset contributes significantly to enhancing tree health monitoring, improving crop yield, and promoting sustainable agricultural practices.

Why it matches plant phenotyping methods植物葉の病害症状を画像で分類する大規模データセットであり、植物の病害状態を直接評価する再利用可能なベンチマーク資源が中心です。

abstractThis study presents a large, meticulously curated and manually validated dataset aimed at classifying leaf quality into five critical categories: Healthy, Bacterial Spot, Shot Hole, Yellow, and Powdery Mildew.
Reproduction assets foundThe paper is a Data in Brief article describing AI-MedLeafX, a public leaf-image dataset for medicinal plant disease classification, deposited on Mendeley Data with an explicit direct URL and DOI (10.17632/zz7r5y4dc6.1). This is the paper's own phenotyping image dataset (10,858 original images, 65,148 augmented, 13疾病/类
Dataset · publicification tasks and future agricultural research applications . Data source location National Botanical Garden, Mirpur-2, Dhaka – 1216 Latitude: 23.8121° N Longitude: 90.3531° E Zone: Dhaka Country: Bangladesh Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/zz7r5y4dc6.1 Direct URL to data: https://data.mendeley.com/datasets/zz7r5y4dc6/1 The dataset is published under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Related research article None 1. Value of the Data • This is a unique and complete dataset of images from different categories, including healthy, bacterial spot, shot hole, powdery mildew, and yellow leaf. This datasetOpen asset ↗Mendeley Data · 10.17632/zz7r5y4dc6.1lines:1-49
Code / dataset availability confirmedCrossref · Europe PMC · checked 14 Sept 2026
Published1 Aug 2025Data in BriefCited by 3 · OpenAlex ↗

TomatoWUR: An annotated dataset of tomato plants to quantitatively evaluate segmentation, skeletonisation, and plant-trait extraction algorithms for 3D plant phenotyping

TomatoLiDAR / point cloudLeafStem / branchWhole plant / canopy / plot / fieldMorphology / geometry measurementSegmentationSkeletonization / topologyArchitecture / morphology / geometryLeaf traits

Plant phenotyping involves the measurements of plant traits to gain more insight into the interaction between the genotype (G), environment (E) and crop management strategies (M). To improve plant phenotyping, accurate measurements are crucial. Manual measurements are biased, time-intensive, and therefore limited to only a few plants. Especially measurements of 3D phenotypic traits, such as plant architecture, internode length, and leaf area are difficult to extract manually. To enhance the speed and accuracy of phenotyping, there is a need for automatic digital plant phenotyping solutions. The presented dataset contains 3D point clouds of tomato plants, which will enable researchers to develop novel methods to extract 3D phenotypic traits. Converting 3D point clouds to plant traits is also known as 3D plant phenotyping. This process can be subdivided into three steps: point cloud segmentation, skeletonisation to extract plant architecture, and plant-traits extraction. Those three steps need to be analysed properly to indicate bottlenecks and improve 3D phenotyping algorithms. Currently, the development of 3D phenotyping algorithms is inhibited by the availability of comprehensive datasets and algorithms to analyse all steps. To our best knowledge only five annotated datasets exist for testing and validating 3D phenotyping algorithms. However, these datasets mainly focus on the segmentation step. Skeletonisation and manual measured plant traits are frequently not included. To improve 3D plant phenotyping, a novel dataset, TomatoWUR, is presented. This comprehensive dataset consists of 44 point clouds of single tomato plants imaged by fifteen cameras to create a point cloud using the shape-from-silhouette methodology. The dataset includes annotated point clouds, skeletons, and manual reference measurements. In addition, the dataset includes software for comprehensive evaluation and comparison of phenotyping methods, which is expected to benefit the development of 3D phenotyping algorithms. The related software can be found our GIT: https://github.com/WUR-ABE/TomatoWUR.

Why it matches plant phenotyping methods3D植物フェノタイピング用の注釈付きデータセットと評価ソフトウェアを提示し、セグメンテーション、骨格化、形質抽出アルゴリズムの開発・検証を直接支援するため、方法論が中心である。

abstractThe presented dataset contains 3D point clouds of tomato plants, which will enable researchers to develop novel methods to extract 3D phenotypic traits.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicIn addition, the dataset includes software for comprehensive evaluation and comparison of phenotyping methods, which is expected to benefit the development of 3D phenotyping algorithms. The related software can be found our GIT: https://github.com/WUR-ABE/TomatoWUROpen asset ↗WUR-ABE/TomatoWURlines:1-45
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 15 Sept 2026
Published1 Aug 2025Research SquareCited by 0 · OpenAlex ↗

FIP 1.0 Soybean data: Insights on soybean growth from eight years of high-throughput image field phenotyping

SoybeanField / plotRGB / grayscaleWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenology

Abstract Soybean growth is determined by the interaction of genetic, environmental, and management factors. In the context of future climate and climate extremes, understanding genotype by environment interaction (GxE) will be crucial for selecting resilient breeding lines and optimizing management practices to minimize stress. As stress periods occur periodically in a season, in depth knowledge, about causing weather variables and differing responses of genotypes over time is required. In field studies, however, the environment is often treated as a static factor, and the specific effects of weather variability on growth remain poorly understood. Here, we present a longitudinal dataset comprising 17,247 high-resolution RGB images of soybean breeding line collected over eight years in Eschikon, Switzerland. Top of canopy images were acquired throughout the entire growing seasons and complemented by hourly weather data, enabling a comprehensive analysis of soybean growth dynamics under varying field conditions. High spatio-temporal image resolution enables detailed analysis of growth dynamics and GxE, supporting identification of stress-tolerant genotypes to improve yield prediction and yield stability.

Why it matches plant phenotyping methods8年間の高解像度RGB画像による圃場フェノタイピングデータセットを提示し、作物生育動態とG×E解析を可能にする方法・データ基盤が中心である。

titleFIP 1.0 Soybean data: Insights on soybean growth from eight years of high-throughput image field phenotyping
Reproduction assets foundThe paper is a data note whose core contribution is a public soybean phenotyping dataset (raw FIP images, segmentation masks, canopy cover data, BLUEs, weather, reference traits) deposited at ETH Research Collection, plus the authors' canopy cover extraction workflow code on GitLab. Both are paper-specific, public, and
Dataset · publicason, therefore, from 2020 to 2022, photosynthetic photon fluence rate (PPFR) was taken from a LI-COR sensor placed next to the field. The factor to convert radiation in MJ m− 2 to PPFR was 2.04 according to [26]. 4.1 Data Files and Structure The dataset presented in this study is available at ETH Research Collection under DOI: https://doi.org/10.3929/ethz-b-000742401. The dataset is structured into directories that align with the described data processing pipeline used for extracting and analyzing canopy cover traits from field images. All files are provided in interoperable and widely-used ‘.csv‘ and ‘.png‘ format. • data/Design 2015 2022 Eschikon.csv: Experimental design file, including pOpen asset ↗ETH Research Collection · 10.3929/ethz-b-000742401pdf-raw-page:6 lines:1-47
Code · publicCollection (https://doi.org/10.3929/ethz-b-000742401) and as Hugging Face data set card (doi.org/10.57967/hf/6052) allowing interoperability and standardization with other datasets. 9 Code availability Users with similar data can use the implemented workflow to get canopy cover from their experiments. The code is available on: https://gitlab.ethz.ch/crop_phenotyping/fip-soybean-canopycover 10 Author contributions BK: Developed algorithm, analyzed data and drafted manuscript; NK, LR, AH, AM: FIP development, BK, NK, CO, LK, LR, OZ, SC, FL, HA, NS, FT, HZ, CAB, CB, AH: Collected and prepared data; Experimental design: BK, LK, LR, AH; all authors improved and approved the manuscript 5Open asset ↗gitlab.ethz.ch · crop_phenotyping/fip-soybean-canopycoverpdf-raw-page:7 lines:1-46
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published31 Jul 2025PloS oneCited by 0 · OpenAlex ↗

CSCA-YOLOv8: A lightweight network model for evaluating drought resistance in mung bean.

Chlorophyll fluorescenceWhole plant / canopy / plot / fieldClassificationStress / disease detectionStress response / tolerance

Drought is one of the main factors affecting mung bean production in China. Screening drought-resistant germplasm resources and cultivating drought-resistant varieties are of great significance to the development of the mung bean industry in China. Combined with chlorophyll fluorescence imaging technology, this paper proposes a lightweight mung bean drought resistance identification network model based on YOLOv8, referred to as CSCA-YOLOv8. The model uses StarNet to replace the backbone network of YOLOv8 to reduce the size of the model. The C2f_Star module is introduced in the neck structure instead of the original C2f module. Then, in order to enhance the network's attention to the key regions in the feature map, the Context Anchor Attention Mechanism (CAA) module is also introduced into the fourth C2f_Star module. Then, a CGBD module is proposed in the neck structure to reconstruct the ordinary convolution to improve the feature extraction ability of the model for small targets. Finally, the SIoU loss function is used to replace CIoU to accelerate the convergence of the model. In the actual data analysis, we used the collected 4808 chlorophyll fluorescence images of the natural mung bean population under drought stress to make the Mungbean Drought Datatset(MDD) and made classification labels for each image according to different drought resistance levels, which were 0, 1, 2, 3, 4 and 5. We also verified the excellent performance and generalization performance of the model using the collected MDD dataset. The final experimental results show that compared with the YOLOv8s baseline model, the number of parameters of our proposed algorithm is reduced by 24%, the floating point number is reduced by 35%, and the accuracy is improved by 2.52%, which supports the deployment on embedded edge devices with limited computing power. Therefore, our proposed algorithm has great potential in the field of drought resistance identification and germplasm selection of mung bean.

Why it matches plant phenotyping methods乾燥抵抗性を推定するクロロフィル蛍光画像ベースのYOLOv8改良モデルを開発し、データセット上で性能・汎化性能を検証しており、表現型取得・抽出手法が中心である。

abstractCombined with chlorophyll fluorescence imaging technology, this paper proposes a lightweight mung bean drought resistance identification network model based on YOLOv8
Reproduction assets foundThe paper's data availability statement explicitly deposits the authors' MDD chlorophyll fluorescence image dataset (4808 mung bean drought-resistance images with labels) and their CSCA-YOLOv8 source code on a public GitHub repository, making both directly actionable paper-specific assets.
Code · publicThe dataset and source code are available on Github.Open asset ↗pdf-page:3 lines:1-51
Code / dataset availability confirmedarXiv · checked 6 Sept 2026
Published15 Jul 2025arXiv

Tomato Multi-Angle Multi-Pose Dataset for Fine-Grained Phenotyping

TomatoRGB / grayscaleFlowerFruitPanicle / ear / spikeLeafStem / branchWhole plant / canopy / plot / fieldClassificationObject detection

Observer bias and inconsistencies in traditional plant phenotyping methods limit the accuracy and reproducibility of fine-grained plant analysis. To overcome these challenges, we developed TomatoMAP, a comprehensive dataset for Solanum lycopersicum using an Internet of Things (IoT) based imaging system with standardized data acquisition protocols. Our dataset contains 64,464 RGB images that capture 12 different plant poses from four camera elevation angles. Each image includes manually annotated bounding boxes for seven regions of interest (ROIs), including leaves, panicle, batch of flowers, batch of fruits, axillary shoot, shoot and whole plant area, along with 50 fine-grained growth stage classifications based on the BBCH scale. Additionally, we provide 3,616 high-resolution image subset with pixel-wise semantic and instance segmentation annotations for fine-grained phenotyping. We validated our dataset using a cascading model deep learning framework combining MobileNetv3 for classification, YOLOv11 for object detection, and MaskRCNN for segmentation. Through AI vs. Human analysis involving five domain experts, we demonstrate that the models trained on our dataset achieve accuracy and speed comparable to the experts. Cohen's Kappa and inter-rater agreement heatmap confirm the reliability of automated fine-grained phenotyping using our approach.

Why it matches plant phenotyping methods植物の多視点画像取得、アノテーション付きデータセット、深層学習による分類・検出・セグメンテーションを中心に開発・検証した、明確な植物フェノタイピング手法研究です。

abstractwe developed TomatoMAP, a comprehensive dataset for Solanum lycopersicum using an Internet of Things (IoT) based imaging system with standardized data acquisition protocols.
Reproduction assets foundThe paper's TomatoMAP dataset (images, annotations) is publicly deposited in e!DAL at IPK with an explicit DOI URL given in the Data Records section.
Dataset · publicDataset is deposited in e!DAL (electronic data archive library) of IPK (Leibniz Institute of Plant Genetics and Crop Plant Research): https://doi.ipk-gatersleben.de/DOI/10bb9f14-ce90-4747-836f-cf61dfb5eea1/Open asset ↗e!DAL · 10bb9f14-ce90-4747-836f-cf61dfb5eea1pdf-page:7 lines:1-73
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published12 Jul 2025Scientific dataCited by 9 · OpenAlex ↗

Variation of winter wheat phenology dataset in Huang Huai Hai Plain of China from 1981 to 2021.

WheatField / plotWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenology

This study presents a comprehensive analysis of winter wheat phenological variations in China's Huang-Huai-Hai Plain (HHHP) from 1981 to 2021, leveraging data from 62 national agrometeorological observation stations. As the world's largest winter wheat production region, the HHHP contributes over 60% of China's total output, playing a pivotal role in national food security. Using kernel density estimation (KDE) and univariate linear regression, the dataset characterizes interannual trends in key phenological stages-sowing, emergence, tillering, jointing, booting, heading, flowering, milking, and maturity-along with growth period durations. Results reveal significant shifts in phenological timings and growth stages under climate change, such as advanced heading stages and altered phase lengths, which correlate with temperature increases and extreme weather events. The dataset, comprising 1,120 figures generated via Origin Lab, is publicly available on ScienceDB, providing critical insights for climate adaptation strategies, cultivation optimization, and yield stability. Technical validation confirms the reliability of the data, sourced from standardized, long-term manual observations by trained professionals under China Meteorological Administration protocols. This work offers a foundational resource for understanding climate-crop interactions and guiding sustainable agricultural practices in a warming world.

Why it matches plant phenotyping methods冬小麦の複数生育ステージという植物形質を長期・標準化観測で収録した公開データセットであり、データの技術的検証も含むため、フェノタイピングデータセットとして中心的です。

abstractthe dataset characterizes interannual trends in key phenological stages-sowing, emergence, tillering, jointing, booting, heading, flowering, milking, and maturity-along with growth period durations
Reproduction assets foundThe paper describes a public dataset of winter wheat phenology (1,120 KDE and linear-trend figures from 62 agrometeorological stations, 1981–2021) deposited on ScienceDB under DOI 10.57760/sciencedb.23011, freely downloadable. No custom analysis code exists ('No custom code was created for the production of this dataet
Dataset · publicThe Variation of winter wheat phenology dataset in Huang Huai Hai Plain of China from 1981 to 2021 is available at ScienceDB 35 . The dataset is provided in JPG format estimated and plotted by Origin Lab. All the diagrams can be downloaded directly for free.Open asset ↗lines:47-83
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Jul 2025Data in briefCited by 1 · OpenAlex ↗

Image dataset of Taro Leaf Blight disease collected from the West African Sub-Region.

TaroField / plotLeafStress / disease detectionDisease symptoms / severity

This dataset encompasses an extensive collection of 18,248 high-resolution JPEG images, documenting various stages of Taro Leaf Blight (TLB) infection in Taro plants across West Africa. TLB, primarily caused by the pathogen Phytophthora colocasiae, manifests through necrotic leaf spots, white sporangia bands, and orange droplets, severely impacting the agricultural output and economic stability of smallholder farmers in the region. The images represent a range of infection stages-early, mid, late, and healthy conditions-captured during the dry and early rainy seasons in Nigeria and Ghana using smartphones equipped with high-resolution cameras. This dataset was carefully curated to help in the development and training of machine learning models for early and accurate detection of TLB, a crucial step towards effective disease management. By enabling the application of advanced diagnostics through technologies such as smartphone apps and AI-based analysis tools, this dataset not only aims to enhance the technological capabilities within agricultural sectors but also serves as a vital educational resource. Researchers and developers can utilize this dataset to create and refine models that diagnose plant diseases promptly, thereby allowing for timely interventions that can prevent widespread crop damage and subsequent economic losses. Additionally, the dataset supports ongoing efforts to integrate artificial intelligence with traditional farming practices, offering a bridge between advanced technological solutions and accessible applications for resource-limited settings. The potential reuse of this dataset extends beyond disease identification; it encompasses agricultural research, educational purposes, and further development of automated systems for plant health monitoring, making it a cornerstone for future innovations in agricultural technology and management strategies.

Why it matches plant phenotyping methodsタロイモ葉の病害症状を画像で記録した大規模データセットであり、植物の病害状態を推定する画像ベース表現型解析の基盤として、データセット自体が中心的成果である。

abstractThis dataset encompasses an extensive collection of 18,248 high-resolution JPEG images, documenting various stages of Taro Leaf Blight (TLB) infection in Taro plants across West Africa.
Reproduction assets foundThe paper is a Data in Brief describing a public plant-phenotyping image dataset (18,248 taro leaf blight images) deposited on Mendeley Data with an explicit direct URL and DOI, matching an allowed URL exactly.
Dataset · publiction • Institution : University of Lagos, Akoka. Kwame Nkrumah University of Science and Technology • City/Town/Region: Abakaliki, Ebonyi, Izzi, Ezza North, Agbani, Ngwo, Ashanti. • Country : Nigeria and Ghana Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/3knm93dkc5.1 Direct URL to data: https://data.mendeley.com/datasets/3knm93dkc5/1 Related research article Nwaneto, C., Yiinka-Banjo, C., Ugot, O. A., Annor, T., & Umeugochukwu, O. (2024). EARLY DETECTION OF THE TARO LEAF BLIGHT DISEASE IN THE WEST AFRICAN SUB-REGION USING DEEP IMAGE CLASSIFICATION MODELS. Smart Agricultural Technology , 100,636. 1 Value of the Data • This dataset is important for devOpen asset ↗Mendeley Data · 10.17632/3knm93dkc5.1lines:1-60
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published9 Jul 2025Data in briefCited by 1 · OpenAlex ↗

Dataset of apples for grading by sweetness, ripeness and variety.

AppleLaboratory / benchtopMultispectral / hyperspectralFruitClassificationGrowth / development / phenology

The study created a detailed database for apple quality inspection using a cost-effective, self-designed multi-spectral imaging system. The system was optimized to allow spectral information to be obtained in 8 discrete wavebands, which enabled non-destructive determination of such key fruit components as ripeness, sugar type and cultivar. Stringent environmental conditions were maintained during image acquisition for optimal measurement consistency and experimental repeatability. The detailed dataset encompasses 32,463 multi-spectral images across three distinct classification categories. For sweetness evaluation, 1620 images spanning Brix values from 10 % to 15 % were collected from five apple varieties. Ripeness evaluation includes 29,160 images documenting the complete maturation cycle over 18 days, while variety classification contains 1683 images from three distinct cultivars. Each image was captured under controlled lighting conditions using eight specific wavelengths, ensuring spectral consistency crucial for machine learning applications. These multi-spectral images were concatenated for grading by sweetness, ripeness, and variety, creating a processed dataset of concatenated images optimized for AppleNet processing. The concatenation process combines the eight wavelength channels into unified image representations suitable for deep learning applications. Sample collection included the picking of different apple cultivars at different physiological development phases of fruit from local orchards. Single specimens were imaged sequentially using a multi-spectral technique. Information on sugar content concentration ( % Brix), maturation phase classification and varietal identification was recorded according to standard laboratory procedure. The resulting annotated database includes such quantitative reference points, which can be used to train supervised learning classifiers in computational classification systems. The reuse value of the dataset covers a wide range of applications such as machine learning-based fruit quality evaluation, agricultural automation and food industry examination. This dataset of ours can be used by researchers to develop and test algorithms to classify apples and estimate their ripeness and the presence of diseases. Furthermore, the proposed multi-spectral imaging can be generalized to cover other fruits and agricultural products, extending the application of the method in smart agriculture. This dataset serves as a valuable resource for researchers in computer vision, machine learning, and agricultural technology, fostering advancements in non-destructive fruit quality evaluation methodologies.

Why it matches plant phenotyping methodsリンゴの甘度・成熟度・品種を推定するマルチスペクトル撮像システムと、注釈付き大規模画像データセットを中心に構築しており、植物器官の品質・状態を定量化する再利用可能なフェノタイピング手法に該当する。

abstractThe study created a detailed database for apple quality inspection using a cost-effective, self-designed multi-spectral imaging system.
Reproduction assets foundThe article is a Data in Brief describing a public multi-spectral apple image dataset (sweetness/Brix, ripeness over 18 days, variety) deposited on Mendeley Data with DOI 10.17632/y5h6v8w6ms.2 and a direct URL, explicitly stated as publicly accessible. This is the paper's own phenotyping image dataset. The MATLAB code,
Dataset · publicme environment using a custom-built multi-spectral imaging chamber . The imaging conditions were carefully maintained to ensure consistency. The dataset is securely stored for research and study purposes. Data accessibility Repository name: Mendeley Data Data identification number: DOI: 10.17632/y5h6v8w6ms.2 Direct URL to data: https://data.mendeley.com/datasets/y5h6v8w6ms/2 Instructions for accessing these data: Dataset Title: Dataset of Apples for Grading by Sweetness, Ripeness, and Variety Public Access: The dataset titled ``Dataset of Apples for Grading by Sweetness, Ripeness, and Variety'' is publicly available on Mendeley Data and can be accessed via the following DOI: https://doi.org/Open asset ↗Mendeley Data · 10.17632/y5h6v8w6ms.2lines:40-82
Code / dataset availability confirmedCrossref · checked 13 Sept 2026
Published8 Jul 2025Mathematics

Detection of Citrus Huanglongbing in Natural Field Conditions Using an Enhanced YOLO11 Framework

CitrusField / plotLeafObject detectionDisease symptoms / severity

Citrus Huanglongbing (HLB) is one of the most devastating diseases in the global citrus industry, but its early detection under complex field conditions remains a major challenge. Existing methods often suffer from insufficient dataset diversity and poor generalization, and struggle to accurately detect subtle early-stage lesions and multiple HLB symptoms in natural backgrounds. To address these issues, we propose an enhanced YOLO11-based framework, DCH-YOLO11. We constructed a multi-symptom HLB leaf dataset (MS-HLBD) containing 9219 annotated images across five classes: Healthy (1862), HLB blotchy mottling (2040), HLB Zinc deficiency (1988), HLB yellowing (1768), and Canker (1561), collected under diverse field conditions. To improve detection performance, the DCH-YOLO11 framework incorporates three novel modules: the C3k2 Dynamic Feature Fusion (C3k2_DFF) module, which enhances early and subtle lesion detection through dynamic feature fusion; the C2PSA Context Anchor Attention (C2PSA_CAA) module, which leverages context anchor attention to strengthen feature extraction in complex vein regions; and the High-efficiency Dynamic Feature Pyramid Network (HDFPN) module, which optimizes multi-scale feature interaction to boost detection accuracy across different object sizes. On the MS-HLBD dataset, DCH-YOLO11 achieved a precision of 91.6%, recall of 87.1%, F1-score of 89.3, and mAP50 of 93.1%, surpassing Faster R-CNN, SSD, RT-DETR, YOLOv7-tiny, YOLOv8n, YOLOv9-tiny, YOLOv10n, YOLO11n, and YOLOv12n by 13.6%, 8.8%, 5.3%, 3.2%, 2.0%, 1.6%, 2.6%, 1.8%, and 1.6% in mAP50, respectively. On a publicly available citrus HLB dataset, DCH-YOLO11 achieved a precision of 82.7%, recall of 81.8%, F1-score of 82.2, and mAP50 of 89.4%, with mAP50 improvements of 8.9%, 4.0%, 3.8%, 3.2%, 4.7%, 3.2%, and 3.4% over RT-DETR, YOLOv7-tiny, YOLOv8n, YOLOv9-tiny, YOLOv10n, YOLO11n, and YOLOv12n, respectively. These results demonstrate that DCH-YOLO11 achieves both state-of-the-art accuracy and excellent generalization, highlighting its strong potential for robust and practical citrus HLB detection in real-world applications.

Why it matches plant phenotyping methods柑橘葉のHLB症状を画像から検出するYOLOベースの表現型取得手法を開発し、専用データセットと公開データセットで性能検証しているため、植物病害表現型の方法研究として中心的である。

abstractwe propose an enhanced YOLO11-based framework, DCH-YOLO11.
Reproduction assets foundThe paper's authors publicly release their DCH-YOLO11 model implementation and analysis code on GitHub, as stated in the Data Availability Statement. The MS-HLBD image dataset itself is not stated as publicly deposited (further materials only by request), so only the code asset qualifies.
Code · publicThe project’s code and model implementation are publicly available at https://github.com/CdW8/DCH-YOLO11 (accessed on 6 July 2025).Open asset ↗CdW8/DCH-YOLO11pdf-page:23 lines:1-59
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published7 Jul 2025Plants (Basel, Switzerland)Cited by 30 · OpenAlex ↗

Resource-Efficient Cotton Network: A Lightweight Deep Learning Framework for Cotton Disease and Pest Classification.

CottonClassificationStress / disease detectionDisease symptoms / severity

Cotton is the most widely cultivated natural fiber crop worldwide, yet it is highly susceptible to various diseases and pests that significantly compromise both yield and quality. To enable rapid and accurate diagnosis of cotton diseases and pests-thus supporting the development of effective control strategies and facilitating genetic breeding research-we propose a lightweight model, the Resource-efficient Cotton Network (RF-Cott-Net), alongside an open-source image dataset, CCDPHD-11, encompassing 11 disease categories. Built upon the MobileViTv2 backbone, RF-Cott-Net integrates an early exit mechanism and quantization-aware training (QAT) to enhance deployment efficiency without sacrificing accuracy. Experimental results on CCDPHD-11 demonstrate that RF-Cott-Net achieves an accuracy of 98.4%, an F1-score of 98.4%, a precision of 98.5%, and a recall of 98.3%. With only 4.9 M parameters, 310 M FLOPs, an inference time of 3.8 ms, and a storage footprint of just 4.8 MB, RF-Cott-Net delivers outstanding accuracy and real-time performance, making it highly suitable for deployment on agricultural edge devices and providing robust support for in-field automated detection of cotton diseases and pests.

Why it matches plant phenotyping methods綿花の病害を画像から分類する軽量深層学習モデルと画像データセットを開発・評価しており、植物の病害状態を抽出するフェノタイピング手法が中心である。

abstractwe propose a lightweight model, the Resource-efficient Cotton Network (RF-Cott-Net), alongside an open-source image dataset, CCDPHD-11, encompassing 11 disease categories.
Reproduction assets foundThe paper's authors publicly released their self-constructed cotton disease/pest image dataset CCDPHD-11 (18,953 images, 11 classes) on GitHub, as stated in the Data Availability Statement. No code or trained model deposit is mentioned.
Dataset · publicew and editing, K.C., H.W., P.W.C. and R.-F.W.; visualization, H.-W.Z. and R.-F.W.; supervision, H.W., P.W.C. and R.-F.W.; project administration, H.W., P.W.C. and R.-F.W. All authors have read and agreed to the published version of the manuscript. Data Availability Statement The proposed CCDPHD-11 dataset can be found online ( https://github.com/SweefongWong/CCDPHD-11-Dataset , accessed on 9 March 2025). Conflicts of Interest The authors declare no conflicts of interest. Funding Statement This research received no external funding. Footnotes Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributOpen asset ↗SweefongWong/CCDPHD-11-Datasetlines:384-405
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 6 Sept 2026
Published6 Jul 2025PlantsCited by 4 · OpenAlex ↗

StomaYOLO: A Lightweight Maize Phenotypic Stomatal Cell Detector Based on Multi-Task Training.

MaizeMicroscopyLeafStomata / guard-cell complexObject detectionStomatal traits

L.), a vital global food crop, relies on its stomatal structure for regulating photosynthesis and responding to drought. Conventional manual stomatal detection methods are inefficient, subjective, and inadequate for high-throughput plant phenotyping research. To address this, we curated a dataset of over 1500 maize leaf epidermal stomata images and developed a novel lightweight detection model, StomaYOLO, tailored for small stomatal targets and subtle features in microscopic images. Leveraging the YOLOv11 framework, StomaYOLO integrates the Small Object Detection layer P2, the dynamic convolution module, and exploits large-scale epidermal cell features to enhance stomatal recognition through auxiliary training. Our model achieved a remarkable 91.8% mean average precision (mAP) and 98.5% precision, surpassing numerous mainstream detection models while maintaining computational efficiency. Ablation and comparative analyses demonstrated that the Small Object Detection layer, dynamic convolutional module, multi-task training, and knowledge distillation strategies substantially enhanced detection performance. Integrating all four strategies yielded a nearly 9% mAP improvement over the baseline model, with computational complexity under 8.4 GFLOPS. Our findings underscore the superior detection capabilities of StomaYOLO compared to existing methods, offering a cost-effective solution that is suitable for practical implementation. This study presents a valuable tool for maize stomatal phenotyping, supporting crop breeding and smart agriculture advancements.

Why it matches plant phenotyping methodsトウモロコシの気孔を画像から検出するモデルとデータセットを開発・評価しており、植物フェノタイピング手法が研究の中心です。

abstractwe curated a dataset of over 1500 maize leaf epidermal stomata images and developed a novel lightweight detection model, StomaYOLO
Reproduction assets foundThe paper's analysis code (StomaYOLO detector) is openly available on GitHub with an authors' URL; the phenotype image dataset itself is only available on request from the corresponding author.
Code · publicThe code that support the findings of this study are openly available in GitHub at https://github.com/yangziqi2003/StomaYOLO (accessed on 5 May 2025).Open asset ↗https://github.com/yangziqi2003/StomaYOLO · StomaYOLOlines:426-440
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 Jul 2025Data in briefCited by 3 · OpenAlex ↗

Okra disease dataset for classification and segmentation: Dataset collection, analysis and applications.

Field / plotLeafClassificationSegmentationDisease symptoms / severity

The early diagnosis of okra leaf diseases is crucial for maintaining crop health and ensuring high agricultural productivity. To facilitate the development of robust deep learning models for automated disease detection, we present a comprehensive dataset of 2500 okra leaf images collected from real-time agricultural fields in India. The dataset consists of six classes, including healthy leaves (Class 0) and five diseased categories: Leaf Curly Virus (Class 1), Alternaria Leaf Spot (Class 2), Cercospora Leaf Spot (Class 3), Phyllosticta Leaf Spot (Class 4), and Downy Mildew (Class 5). Each image is resized to 224 × 224 pixels to ensure compatibility with standard deep learning models. The primary objective of this dataset collection is to provide a benchmark resource for researchers working on early-stage plant disease classification, detection and segmentation. This dataset is unique as it is one of the first publicly available Indian okra leaf disease datasets captured in real-world conditions, incorporating natural variations in lighting, leaf positioning, and environmental factors. It serves as a valuable resource for future young researchers in the field of smart agriculture, enabling advancements in machine learning-based disease diagnosis, smart farming applications, and precision agriculture. Future enhancements will focus on expanding the dataset with more images, including different growth stages and environmental conditions, to improve model generalization and real-world applicability.

Why it matches plant phenotyping methods植物葉の病害状態を画像で分類・セグメンテーションする公開データセットを構築し、ベンチマーク資源として提供することが中心であるため、植物フェノタイピング手法文献に含める。

abstractwe present a comprehensive dataset of 2500 okra leaf images collected from real-time agricultural fields in India.
Reproduction assets foundThe paper is a Data in Brief article presenting the authors' own Okra DiseaseNet dataset of 2500 okra leaf images for disease classification and segmentation, with an explicit public Mendeley Data deposit (DOI 10.17632/nh7zk4hv8z.1) matching an allowed URL. Other listed URLs are cited prior-work datasets, not paper-own
Dataset · publicr district (PIN: 613006), with latitude and longitude coordinates available at Google Maps link ). These locations were selected to ensure diverse environmental conditions for the dataset collection. Data accessibility Dataset Name: Okra DiseaseNet Dataset Data Identification Number: DOI: 10.17632/nh7zk4hv8z.1 Direct URL link : https://data.mendeley.com/datasets/nh7zk4hv8z/1 Related research article None 1 Value of the Data The Okra Leaf Disease Dataset is the first dataset collected from Indian agricultural farmlands, specifically for okra crop disease analysis. While many plant disease datasets exist, this dataset stands out due to its high-resolution images (enabled by superior camera lenOpen asset ↗10.17632/nh7zk4hv8z.1lines:40-68
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published30 Jun 2025Data in briefCited by 5 · OpenAlex ↗

A durian leaf image dataset of common diseases in Vietnam for agricultural diagnosis.

Field / plotLeafClassificationDisease symptoms / severity

Agriculture plays a vital role in Vietnam's economy, with durian being a key high-value crop that supports millions of farmers. However, durian leaves are highly susceptible to pests, diseases, and environmental stressors, negatively impacting yield and quality. This study introduces a dataset of 2595 durian leaf images, categorized into six classes: 484 healthy leaves and 2111 diseased leaves spanning Blight (440), Colletotrichum (400), Algal (462), Phomopsis (411), and Rhizoctonia (398). The images were collected from durian orchards across Vietnam under diverse conditions, then background-removed, resized to 400 × 400 pixels, and manually annotated with expert guidance. This dataset provides a valuable resource for advancing research in automated plant disease detection, enabling the development of computer vision models for early diagnosis and precision farming, thereby supporting sustainable durian production and improved crop productivity.

Why it matches plant phenotyping methods罹病葉の画像を収集・前処理・専門家注釈したデータセットの構築が中心で、植物病害状態の画像ベース表現型判定に直接利用できる。

abstractThis study introduces a dataset of 2595 durian leaf images, categorized into six classes
Reproduction assets foundThe paper is a data descriptor for a durian leaf disease image dataset (2595 annotated images) publicly deposited on Mendeley Data with DOI 10.17632/pxzvksbwnj and direct URL provided in the text.
Dataset · publiclocations mentioned below. Location 1: Bu Dang District, Binh Phuoc Province, Vietnam Location 2: Cai Lay District, Tien Giang Province, Vietnam Data accessibility Repository: A Durian Leaf Image Dataset of Common Diseases in Vietnam for Agricultural Diagnosis Data identification number: 10.17632/pxzvksbwnj Direct URL to data: https://data.mendeley.com/datasets/pxzvksbwnj Instructions for accessing these data: The data is divided into three sets: train, test, and validation. Each set contains six folders with pre-processed images in JPG format. Related research article 1. Value of the Data • The dataset consists of 2595 images of durian leaves captured via mobile devices, including healthy lOpen asset ↗10.17632/pxzvksbwnjlines:1-54
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published18 Jun 2025The New phytologistCited by 4 · OpenAlex ↗

Automated extraction of leaf mass per area from digitized herbarium specimens.

LeafMorphology / geometry measurementLeaf traits

The digitization of vast herbarium collections has made millions of plant specimen images freely available online, which can now be used to generate phenotypic datasets of unprecedented scope. Here, we assess the potential of computer vision tools to automate the extraction of predicted leaf mass per area (LMA pred ) from digitized herbarium specimens. We use an automated pipeline to extract leaf area and petiole width from 22 680 leaves, representing a phylogenetic informed sample of 1580 species of woody angiosperms. LMA pred is estimated using a proxy equation that models the scaling relationship between petiole width and leaf mass. We assess potential sources of error in LMA pred estimates and evaluate whether documented LMA-climate patterns are recovered using this dataset and phylogenetic comparative methods. Our LMA pred dataset responds mainly to temperature and solar radiation and presents a positive correlation with latitude. The proxy equation, not the automated pipeline, is responsible for most of the error in LMA pred estimates. Our pipeline underscores the power of combining herbarium digitization with new techniques for automated trait scoring. The increased size of datasets generated using this tool allows investigation of potential LMA-climate relationships with a geographically balanced sample while also utilizing comprehensive phylogenetic information.

Why it matches plant phenotyping methodsデジタル標本画像から葉面積・葉柄幅を自動抽出し、LMAを推定するコンピュータビジョン・パイプラインが中心であり、植物形質データセットの生成と誤差評価も行っている。

abstractHere, we assess the potential of computer vision tools to automate the extraction of predicted leaf mass per area (LMA pred ) from digitized herbarium specimens.
Reproduction assets foundThe paper's Data Availability Statement explicitly deposits all data and code (the leaf_vision pipeline for automated LMA extraction from herbarium images) at the authors' public GitHub repository, which is an allowed URL. Supporting Information also contains the full LM2 measurement results (Table S2).
Code · publicAll data and code used here are available on https://github.com/tncvasconcelos/leaf_vision and in the Supporting Information . GBIF DOI is available in the reference list.Open asset ↗tncvasconcelos/leaf_visionlines:358-360
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published16 Jun 2025Plant phenomics (Washington, D.C.)Cited by 4 · OpenAlex ↗

Generation of labeled leaf point clouds for plants trait estimation.

LiDAR / point cloudLeafMorphology / geometry measurement2D/3D reconstructionSkeletonization / topologyLeaf traits

Today, leaf trait estimation remains a labor-intensive process. The effort to obtain ground truth measurements limits how accurately this task can be performed automatically. Traditionally, plant scientists manually measure the traits of harvested leaves and associate them with sensor data, which is key for training machine learning approaches and to automate the processes. In this paper, we propose a neural network-based method to generate synthetic 3D point clouds of leaves with their associated traits to support approaches for phenotyping. We use real-world leaf point clouds to learn how to generate realistic leaves from a leaf skeleton, which is automatically extracted. We use the generated leaves to fine-tune different leaf trait estimation methods. We evaluate our generated data using different trait estimation methods and compare the results to using real-world data or other synthetic datasets from agricultural simulation software. Experiments show that our approach generates leaf point clouds with high similarity to real-world leaves. Tuning trait estimation methods on our generated data improves their performance in the estimation of real-world leaves' traits, making our data crucial for developing and testing data-driven trait estimation methods. Accurate trait estimation is key to understanding crop growth, productivity, and pest resistance, as leaf size directly influences photosynthesis, yield potential, and vulnerability to insects and fungal growth.

Why it matches plant phenotyping methods葉の形質推定を支援するため、形質付き合成3D点群を生成するニューラルネットワーク手法とデータセットを開発・評価しており、フェノタイピング手法が中心である。

abstractwe propose a neural network-based method to generate synthetic 3D point clouds of leaves with their associated traits to support approaches for phenotyping.
Reproduction assets foundThe paper uses two public 3D plant point-cloud datasets (Pheno4D and BonnBeetClouds3D) as real-world inputs for training/evaluating its leaf trait estimation and generation pipeline; both have explicit public URLs. The authors' code is only promised ('We plan to make our code publicly available'), so it is not yet an a
Dataset · publicWe use two publicly available datasets. Pheno4D [43] is available at the url: https://www.ipb.uni-bonn.de/data/pheno4d/index.html. It contains maize and tomato plants measured daily, over 12 and 20 days respectively.Open asset ↗Pheno4Dhtml-lines:401-424
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Jun 2025Data in briefCited by 18 · OpenAlex ↗

teaLeafBD: A comprehensive image dataset to classify the diseased tea leaf to automate the leaf selection process in Bangladesh.

TeaLeafClassificationStress / disease detectionDisease symptoms / severity

Tea is an extremely popular beverage around the world due to its exquisite taste and flavor. Unfortunately, it is prone to different types of illness, which can reduce the amount of harvest along with its standard. Among these, leaf infections are a serious concern since they negatively affect the quality of tea leaves. As a consequence, tea producers often encounter a great deal of obstacles and financial losses. Keeping this in mind, a thorough dataset has been compiled, which contains 5278 images of diseased and healthy leaves. The purpose of this dataset is to improve our knowledge of how these conditions impact cultivating tea plants and tea production. These images are collected from a variety of locations and meteorological circumstances, which provide an extensive knowledge of the disease patterns unique to tea leaves. The pictures have been captured with the help of some high-quality devices from different angles and in high resolution to ensure the standard and increase the usability of the dataset. Rigorous steps were followed when preparing the dataset that would be of great help in building a precise artificial intelligence model. The dataset carefully determined and classified six tea leaf diseases: Tea algal leaf spot, Brown Blight, Gray Blight, Helopeltis, Red spider, and Green mirid bug. There is one more class in the dataset containing images of healthy leaves. These illnesses are known for their devastating impact on tea leaves. An automated disease classification system can be made utilizing deep learning techniques that will enable estate managers to take timely action to stop the spread of the disease, and this meticulously collected dataset will immensely help to train that model.

Why it matches plant phenotyping methods茶葉の健全・病害状態を画像で収集・分類した再利用可能なデータセットであり、植物病害状態の画像ベース表現型計測を中心とする。

abstracta thorough dataset has been compiled, which contains 5278 images of diseased and healthy leaves
Reproduction assets foundThe paper's own teaLeafBD image dataset (5278 tea leaf images, 7 classes) is publicly deposited on Mendeley Data with an explicit DOI and direct URL, making it a paper-specific, publicly actionable asset.
Dataset · public24.30795976, longitude: 91.73760171), 7. Finlay Tea Estate in Sreemangal (latitude: 24.30357177, longitude: 91.74245382), 8. Jungle Bari Tea Estate in Sreemangal (latitude: 24.25329433, longitude: 91.77409053) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/744vznw5k2.4 Direct URL to data: https://data.mendeley.com/datasets/744vznw5k2/4 Related research article None 1. Value of the Data • The dataset was collected from tea harvesting areas in Bangladesh, which is one of the top tea-producing regions in the world, supplying tea globally. A total of 5278 images were captured by the camera from eight tea gardens, and they were annotated by human experts. •Open asset ↗Mendeley Data · 10.17632/744vznw5k2.4lines:1-60
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published4 Jun 2025Cited by 0 · OpenAlex ↗

Multi-Class Banana Leaf Disease Detection via KHO-YOLOv8

Banana / plantainLeafClassificationStress / disease detectionDisease symptoms / severity

Abstract Plant diseases initiate major agricultural challenges because they lead to 16\% of worldwide crop loss in farms. The vulnerability of bananas to diseases including Xanthomonas Wilt and Sigatoka leaf spot puts severe threats to food security because of their extensive damage potential. These diseases possess the risk of damage to the complete harvest so their impact can reach 100%. While deep learning models, particularly YOLO-based architecture, have demonstrated success in plant disease identification, key research gaps remain. One major challenge is the lack of large-scale, annotated datasets for banana leaf diseases, limiting the development and evaluation of robust AI models. Addressing these challenges is crucial, and this study aims to do so by creating a large dataset and developing a robust disease detection model. The dataset comprises more than 5000 samples categorized into three classes: Healthy, Xanthomonas Wilt infected, and Sigatoka leaf spot infected. This study examines a novel framework by employing advanced optimization techniques such as Krill Herd Optimization (KHO) for YOLOv8 and its variants. Our research findings highlight the exceptional performance of the KHO-YOLOv8 model, achieving an impressive accuracy of 96.47%.

Why it matches plant phenotyping methodsバナナ葉の病害状態を画像から分類するデータセット構築とYOLOv8検出モデル開発・評価が研究の中心であり、植物の病徴を直接推定するため対象範囲に該当する。

abstractthis study aims to do so by creating a large dataset and developing a robust disease detection model.
Reproduction assets foundThe paper states its banana leaf disease dataset (5,000 annotated images) and code are publicly available on GitHub, but no concrete repository URL or identifier is provided in the supplied text, and the only allowed URL is an unrelated cited reference. The claim of public availability is explicit, but the asset is not
Dataset · publicMore than 5000 banana leaf images were acquired from different areas of Arba Minch Zuria. Five plant pathologists meticulously verified the image classifications twice to maintain accuracy. Code and dataset is publicly available on GitHub.Open asset ↗GitHubpdf-page:3 lines:1-44
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published3 Jun 2025Copernicus GmbHCited by 1 · OpenAlex ↗

Global near real-time 500 m 10-day FPAR dataset from MODIS and VIIRS for operational agricultural monitoring and crop yield forecasting

Whole plant / canopy / plot / fieldCalibration / preprocessingGrowth / time-series analysisPhotosynthesis / fluorescence

Abstract. Climate change and extreme weather events pose challenges to food security, emphasizing the need for reliable and timely monitoring of crop and rangeland conditions. For this purpose, long-term consistent Earth Observation datasets on vegetation conditions are typically used in early warning and crop yield forecast systems. However, the near-real-time (NRT) production of high quality datasets and the need to guarantee long-term records present various challenges. To address these, we present a NRT global dataset of Fraction of Photosynthetically Active Radiation (FPAR) at 500 m resolution, optimized for agricultural applications. Our dataset combines MODIS-FPAR (Collection 6.1) and VIIRS-FPAR (Collection 2) data, ensuring continuity from 2000 to well beyond 2030. We applied a robust filtering approach based on the Whittaker smoother to produce reliable FPAR estimates in NRT, accounting for sparse and irregular spaced observations due to cloud cover. The dataset is composed of two 10-day filtered timeseries: 1) MODIS-FPAR for 2000 to 2023, being the reference dataset, and 2) intercalibrated VIIRS-FPAR for 2018 onward. While several methods can effectively smooth and gap-fill FPAR data (i.e., using observations before and after the estimation date), our method is designed for optimal filtering in NRT (i.e., using only prior observations). Our approach yields six successive estimates of the same FPAR data point with increasing quality: a inital estimate immediately after the 10-day reference period, four subsequent estimates every 10 days using new observations, and a final consolidated estimate 90 days later. The implemented filtering ingests the available FPAR observations and their original quality assessment (QA) layers. To avoid unrealistic extrapolation when observations are sparse, we impose constraints, season and location specific, to FPAR estimates. We then intercalibrated the VIIRS-FPAR with the MODIS-FPAR filtered timeseries, using a mean difference correction approach, to ensure consistency between both series. This paper describes the filtering and intercalibration method used, the quality assessment of resulting timeseries, and details the obtained products and the corresponding QA layers. The NRT FPAR dataset is publicly available through the Joint Research Centre Data Catalogue, https://data.jrc.ec.europa.eu/dataset/1aac79d8-0d68-4f1c-a40f-b6e362264e50 (Seguini et al., 2025).

Why it matches plant phenotyping methodsFPARという作物・植生キャノピーの明示的な状態量を対象に、MODIS/VIIRSデータのNRTフィルタリング、相互校正、品質評価を開発・記述しており、単なる農業利用ではなく再利用可能な測定データセットと抽出手法が中心である。

abstractwe present a NRT global dataset of Fraction of Photosynthetically Active Radiation (FPAR) at 500 m resolution, optimized for agricultural applications.
Reproduction assets foundThe paper's own filtered and intercalibrated MODIS/VIIRS FPAR dataset (the paper's core output) is explicitly described as open and freely available in near real time via the JRC Data Catalogue and the ASAP website, both of which appear in allowed_urls. No author analysis code is mentioned.
Dataset · publicec.europa.eu/, last access: 30 September 2025) early warning system. The FPAR dataset is accompanied by associated quality layers and has a temporal resolution of 10 d, a time step often used in operational agricultural monitoring. The dataset is open and freely available in NRT through the Joint Research Centre Data Catalogue (https://data.jrc.ec.europa.eu/dataset/1aac79d8-0d68-4f1c-a40f-b6e362264e50, last access: 30 September 2025) and on the ASAP website (https://agricultural-production-hotspots.ec.europa.eu/data/MO6_FPAR, last access: 30 September 2025). This paper has the following specific objectives: (i) to introduce the method used to produce a long-term archive of NRT filtered FPAR Open asset ↗1aac79d8-0d68-4f1c-a40f-b6e362264e50pdf-raw-page:3 lines:1-86
Dataset · publichas a temporal resolution of 10 d, a time step often used in operational agricultural monitoring. The dataset is open and freely available in NRT through the Joint Research Centre Data Catalogue (https://data.jrc.ec.europa.eu/dataset/1aac79d8-0d68-4f1c-a40f-b6e362264e50, last access: 30 September 2025) and on the ASAP website (https://agricultural-production-hotspots.ec.europa.eu/data/MO6_FPAR, last access: 30 September 2025). This paper has the following specific objectives: (i) to introduce the method used to produce a long-term archive of NRT filtered FPAR data; (ii) to present the intercalibration performed between the filtered MODIS-FPAR and the filtered VIIRS- FPAR; (iii) to evaluate tOpen asset ↗ASAP websitepdf-raw-page:3 lines:1-86
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published1 Jun 2025Plant PhenomicsCited by 5 · OpenAlex ↗

XFruitSeg-A general plant fruit segmentation model based on CT imaging.

CitrusX-ray / CTFruitTissueSegmentation

Identification of the phenotypes of fruits is critical for understanding complex genetic traits. Computed tomography (CT) imaging technology enables the noninvasive acquisition of three-dimensional images of fruit interiors, thus providing a robust data foundation for phenotypic analysis. Accurate segmentation of internal fruit tissues is essential, as it directly influences the accuracy and reliability of the results. Current methods are not optimized for the unique features of plant fruit images. This study introduces XFruitSeg, which is a general deep learning model for segmenting plant fruit CT images. The model uses a U-shaped encoder-decoder architecture and integrates multitask learning. A large convolutional kernel network, RepLKNet, expands the receptive field for feature extraction. Multiscale skip connections and a deep supervision mechanism improve the model's capacity to learn features of various sizes, and a contour feature learning branch specifically targets the interorganizational boundaries. An optimized composite loss function enhances the model's robustness when applied to imbalanced categories. Additionally, a dataset named XrayFruitData was established, which contains high-resolution images of twelve plant fruit varieties, with accurate annotations for orange, mangosteen, and durian fruits for model evaluation. Compared with four mainstream advanced models, XFruitSeg achieved superior segmentation performance on the orange, mangosteen, and durian datasets, with mean Dice coefficients of 95.21 ​%, 93.24 ​%, and 94.70 ​% and mean intersection over union (mIoU) scores of 91.09 ​%, 87.91 ​%, and 90.35 ​%, respectively. The results of extensive ablation experiments demonstrate the effectiveness of each component. Therefore, the proposed XFruitSeg model has been proven to be beneficial for high-precision analysis of internal fruit phenotyping traits.

Why it matches plant phenotyping methods果実CT画像から内部組織を分割し、表現型解析を可能にする深層学習モデルと評価用データセットを開発・検証しており、植物フェノタイピング手法が中心である。

abstractThis study introduces XFruitSeg, which is a general deep learning model for segmenting plant fruit CT images.
Reproduction assets foundThe paper's CT fruit segmentation dataset (XrayFruitData), model weights, and source code are publicly available on the authors' GitHub repository, explicitly stated in the Data availability section and dataset description.
Code · publicSome of the raw data, model weights and source codes are accessible at https://github.com/BME-PhenoTeam/Xray4Plant-FruitOpen asset ↗BME-PhenoTeam/Xray4Plant-Fruitlines:530-585
Code / dataset availability confirmedCrossref · Europe PMC · checked 14 Sept 2026
Published1 Jun 2025Data in BriefCited by 12 · OpenAlex ↗

An Indian UAV and leaf image dataset for integrated crop health assessment of soybean crop.

SoybeanAerial / UAVField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Soybean is an important oilseed crop, rich in protein and oil, often referred to as a ``cash crop'' or ``gold bean'' by Indian farmers. In Maharashtra, soybean cultivation spans over approximately 3.8 million hectares, producing 3.07 million tons, placing the state second in India for overall soybean production. However, despite of its significance, several issues such as weeds, diseases, and pests hamper the overall productivity of soybean. Addressing these challenges faced by soybean growers it is essential to enhance yield and improve the crop's overall potential Currently, the farming sector is transitioning towards Agriculture 5.0, also known as digital farming. This approach utilizes data-driven technologies, such as artificial intelligence and computer vision, to transform the agriculture sector. These technologies enable the automation of several farming tasks. To develop accurate and robust machine learning/deep learning models high quality datasets are needed. With this aim, we have created a comprehensive dataset of soybean crop images affected by diseases and pest attacks from original fields of Maharashtra region located in India. Data acquisition was conducted across two seasons through aerial as well as ground-based approaches. The dataset is enriched with 4 types of diseases and 1 pest attack. The proposed dataset will serve as a valuable resource for training and testing machine learning and deep learning models ,enabling accurate detection and classification of diseases and pests attack damage.

Why it matches plant phenotyping methods大豆の病害・害虫被害を対象とした航空・地上画像データセットの構築が中心で、植物の病害状態を画像から評価する再利用可能な資源であるため。

abstractwe have created a comprehensive dataset of soybean crop images affected by diseases and pest attacks from original fields of Maharashtra region located in India.
Reproduction assets foundThe paper's own soybean UAV and leaf image dataset is publicly deposited on Mendeley Data with an explicit direct URL and DOI, matching an allowed URL.
Dataset · publicarashtra, India Banawadi: Longitude 74.1943023 Latitude:17.3179823 Goware: Longitude 74.208238 Latitude:17.286995 Data accessibility Repository name: Mendeley Data MH-SoyaHealthVision: An Indian UAV and Leaf Image Dataset for Integrated Crop Health Assessment Data identification number: 10.17632/hkbgh5s3b7.1 Direct URL to data: https://data.mendeley.com/datasets/hkbgh5s3b7/1 1 Value of the Data • The dataset uniquely combines UAV-based aerial images, offering high resolution and a broad spectrum, with ground-level close-up images. UAV imaging is effective for macro level field variability while ground-based images provide micro level finer details of symptoms such as leaf spots, lesions, andOpen asset ↗Mendeley Data · 10.17632/hkbgh5s3b7.1lines:1-55
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 May 2025Data in briefCited by 15 · OpenAlex ↗

Grapes leaf disease dataset for precision agriculture.

GrapevineField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Grapes are widely cultivated fruit crops, essential for fresh consumption, winemaking and dried product production. However, their yield and quality are significantly impacted by various fungal diseases. This paper provides a large dataset of 2,726 high-quality grape leaf disease images collected from grapes farm of Nashik, India in two years of span 2023 to 2025. The dataset is precisely annotated under the guidance and observation of agriculture domain expert and organized in a well-defined folder structure. The dataset captures the two major categories healthy leaves and unhealthy leaves, during cultivation period. A primary directory containing two main classes Heathy Leaf Images and Unhealthy Leaf images. Further unhealthy class is divided into three subfolders for disease class, namely Downy Mildew, Powdery Mildew and Bacterial Leaf Spot. These are the major fungal disease observed on grape crop causes substantially crop losses and ultimately impact on the yield production. Timely identification of these diseases can significantly reduce the risk of crop loss and help to improve quality of fruit with maximum yield production. This High-quality annotated image dataset can help to design standard advanced AI models for automated disease detection, classification, and prediction. The dataset was validated through a transfer learning approach using the ResNet-18 algorithm and demonstrated the remarkable classification accuracy of 96 % . These results validate the dataset's quality and its suitability for deep learning-based grape disease detection. Overall, this open-access resource provides a valuable foundation for computer vision, machine learning, and agricultural technology researchers aims to enhance disease management practices in grape production. thus, this is an effective source of data for future studies and real-world applications in sustainable grape production.

Why it matches plant phenotyping methodsブドウ葉の病徴・健全状態を画像で取得した注釈付きデータセットを提供し、分類モデルで検証しているため、植物病害表現型のデータセット開発・検証が中心です。

abstractThis paper provides a large dataset of 2,726 high-quality grape leaf disease images collected from grapes farm of Nashik, India in two years of span 2023 to 2025.
Reproduction assets foundThe paper is a data descriptor for the Niphad Grape Leaf Disease Dataset (NGLD), 2,726 annotated grape leaf images, publicly deposited on Mendeley Data with direct URL and DOI. No author analysis code is shared.
Dataset · publicges were labelled sequentially for clear association within the dataset. Data source location Niphad Grapes farms, located at District Nashik 422209, MH-India Longitude and Latitude: 20.0771° N, 74.1094° E Data accessibility Repository Name: Niphad Grape Leaf Disease Dataset (NGLD) DOI: 10.17632/8nnd2ypcv3.5 Direct URL to Data: https://data.mendeley.com/datasets/8nnd2ypcv3/5 1. Value of the Data • Comprehensive Dataset : The Dataset is comprehensive and consists of 2726 high-quality images, in four subfolder such as Downy Mildew, Powdery Mildew, Bacterial Leaf Spot and Healthy Grapes Leaf. Unlike existing public datasets that primarily focus on diseases such as Esca, Black Rot, and Leaf BligOpen asset ↗10.17632/8nnd2ypcv3.5lines:1-43
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 May 2025Data in briefCited by 3 · OpenAlex ↗

A dataset for vineyard disease detection via multispectral imaging.

GrapevineField / plotMultispectral / hyperspectralLeafStem / branchStress / disease detectionDisease symptoms / severity

The present dataset is a collection of multispectral images designed for development of detection algorithms for grapevine diseases like Flavescence dorée (FD) and Esca (ED). Although FD severely threatens viticulture, there are few public datasets and none with multispectral data collected in the field. The collected images have been taken from a frontal perspective of vineyard plants that highlights details of leaves and trunks facilitating detailed disease analysis. The data were collected using a Micasense RedEdge-P multispectral camera, capturing six spectral bands across 172 image captures of three different grapevine varieties used in Lambrusco wines: Ancellotta, Marani, and Salamino. The dataset includes raw and processed images, calibration images for the multispectral camera, annotations detailing plant health conditions, and Python-based usage examples for researchers. Potential applications include the development of machine learning algorithms for automated disease detection, image alignment techniques, and background removal methods. The dataset is a valuable resource for advancing remote and proximal sensing in precision agriculture.

Why it matches plant phenotyping methodsブドウ樹の病害状態を対象とするマルチスペクトル画像データセットで、画像・校正・アノテーション・利用例を含む再利用可能な資源として構築されており、植物フェノタイピング手法の基盤が中心です。

titleA dataset for vineyard disease detection via multispectral imaging.
Reproduction assets foundThe paper is a data descriptor for a multispectral vineyard disease detection dataset deposited by the authors on Zenodo, including raw/processed images, annotations, and Python usage examples. The two Micasense GitHub repositories are generic vendor libraries, not paper-specific assets.
Dataset · publicgio Emilia, Emilia-Romagna, Italy). It is managed by the RIMLab laboratory at the University of Parma, Parco Area delle Scienze 181/A, 43100 Parma, Italy. Data accessibility Repository name: A Dataset for Vineyard Disease Detection via Multispectral Imaging Data identification number: 10.5281/zenodo.14936376 Direct URL to data: https://zenodo.org/records/14936376 Related research article 1 Value of the Data • The dataset features grapevine images which is a high-value plant used for wine production. Italy and other European nations are among the world's largest wine exporters. For this reason, diseases such as Flavescence Dorée (FD) and Esca, that cause severe damage to both the plant aOpen asset ↗Zenodo · 10.5281/zenodo.14936376lines:1-49
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published26 May 2025Frontiers in plant scienceCited by 5 · OpenAlex ↗

Study on the germination rate of maize seeds based on improved YOLOv8n model.

MaizeSeed / grainObject detectionGrowth / development / phenology

The germination potential of corn seeds, a key index for assessing their quality and directly associated with the ultimate corn yield, is currently defined in a way that cannot effectively portray the seed germination rate, and the prevalent measurement methods are traditional, consuming substantial process resources. To tackle these issues, this paper employs a public corn seed germination dataset, adds noise to it to simulate real - world production conditions, and ultimately acquires a dataset comprising 8148 images. It then proposes an enhanced YOLOv8 target detection model, EBS - YOLOv8, for detecting corn seed germination. Specifically, the ECA lightweight attention mechanism is introduced to decrease small - target feature loss, assist in accurate target recognition, and remove redundant features; simultaneously, the P2BiFPN multiscale feature fusion technique is utilized to boost the detection ability for small targets; furthermore, the ScConv convolution is adopted to enhance the feature - extraction capacity and improve detection accuracy. Combined with the improved model, this paper also proposed a mathematical modeling algorithmnew method for measuring seed germination potential and observing seed germination rate. The results indicate that the proposed model attains a mean average precision at 50% Intersection over Union (mAP50) value of 98.9%, a mean average precision in the range of 50% - 95% Intersection over Union (mAP50 - 95) value of 95.8%, an accuracy of 96.7%, and a recall of 96.3%. In comparison with the original model, the mAP50 has increased by 0.9% and the mAP50 - 95 value has witnessed a 3.7% increment. The experiments have demonstrated that the research method for germination potential put forward in this paper can effectively depict the rate variation of seeds during the germination process, thus offering a novel perspective for future research on seed germination potential.

Why it matches plant phenotyping methodsトウモロコシ種子の発芽状態・発芽率を画像から検出・定量する改良YOLOv8モデルと数理測定法を開発し、データセット上で性能検証しているため、植物フェノタイピング手法が中心である。

abstractIt then proposes an enhanced YOLOv8 target detection model, EBS - YOLOv8, for detecting corn seed germination.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicThe dataset used in this experiment can be accessed at ‘ http://dx.doi.org/10.17332/4wkt6thgp6.2 ’.Open asset ↗10.17332/4wkt6thgp6.2lines:333-347
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published26 May 2025Plant methodsCited by 13 · OpenAlex ↗

An advanced deep learning method for pepper diseases and pests detection.

Pepper / chilliGreenhouseWhole plant / canopy / plot / fieldObject detectionStress / disease detectionDisease symptoms / severity

Despite the significant progress in deep learning-based object detection, existing models struggle to perform optimally in complex agricultural environments. To address these challenges, this study introduces YOLO-Pepper, an enhanced model designed specifically for greenhouse pepper disease and pest detection, overcoming three key obstacles: small target recognition, multi-scale feature extraction under occlusion, and real-time processing demands. Built upon YOLOv10n, YOLO-Pepper incorporates four major innovations: (1) an Adaptive Multi-Scale Feature Extraction (AMSFE) module that improves feature capture through multi-branch convolutions; (2) a Dynamic Feature Pyramid Network (DFPN) enabling context-aware feature fusion; (3) a specialized Small Detection Head (SDH) tailored for minute targets; and (4) an Inner-CIoU loss function that enhances localization accuracy by 18% compared to standard CIoU. Evaluated on a diverse dataset of 8046 annotated images, YOLO-Pepper achieves state-of-the-art performance, with 94.26% mAP@0.5 at 115.26 FPS, marking an 11.88 percentage point improvement over YOLOv10n (82.38% mAP@0.5) while maintaining a lightweight structure (2.51 M parameters, 5.15 MB model size) optimized for edge deployment. Comparative experiments highlight YOLO-Pepper's superiority over nine benchmark models, particularly in detecting small and occluded targets. By addressing computational inefficiencies and refining small object detection capabilities, YOLO-Pepper provides robust technical support for intelligent agricultural monitoring systems, making it a highly effective tool for early disease detection and integrated pest management in commercial greenhouse operations.

Why it matches plant phenotyping methodsコショウの病害を画像から検出する深層学習手法を開発し、注釈画像データセットと複数モデルで性能比較・検証しており、植物の病害状態の取得方法が中心です。害虫検出も含みますが、病害検出の技術的貢献が明確なため採用します。

abstractthis study introduces YOLO-Pepper, an enhanced model designed specifically for greenhouse pepper disease and pest detection
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicThe data can be accessed at https://data.mendeley.com/datasets/ disease? Adv Multimed. 2018;2018:6710865.Open asset ↗pdf-page:17 lines:1-54
Code / dataset availability confirmedOpenAlex · checked 15 Sept 2026
Published21 May 2025Cited by 2 · OpenAlex ↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity

Field / plotMultispectral / hyperspectralLeafVisualization / data managementLeaf traitsPhotosynthesis / fluorescence

Abstract. Accurate assessment of leaf functional traits is crucial for a diverse range of applications from crop phenotyping to parameterizing global climate models. Leaf reflectance spectroscopy offers a promising avenue to advance ecological and of robust hyperspectral models for predicting leaf photosynthetic capacity and associated traits from reflectance data has been hindered by limited data availability across species and environments. Here we introduce the Global Spectra-Trait Initiative (GSTI), a collaborative repository of paired leaf hyperspectral and gas exchange measurements from diverse ecosystems. The GSTI repository currently encompasses over 7500 observations from 397 species and 41 sites gathered from 36 published and unpublished studies, thereby offering a key resource for developing and validating hyperspectral models of leaf photosynthetic agricultural research by complementing traditional, time-consuming gas exchange measurements. However, the development capacity. The GSTI database is developed on GitHub (https://github.com/plantphys/gsti) and published to ESS-dive https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2530733, Lamour et al., 2025). It includes gas exchange data, derived photosynthetic parameters, and key leaf traits often associated with traditional gas exchange measurements such as leaf mass per area and leaf elemental composition. By providing a standardized repository for data sharing and analysis, we present a critical step towards creating hyperspectral models for predicting photosynthetic traits and associated leaf traits for terrestrial plants.

Why it matches plant phenotyping methods葉のハイパースペクトル計測とガス交換による光合成形質を結合したデータベースで、植物形質推定モデルの開発・検証を主目的とするため、フェノタイピング手法・データセットとして中心的です。

abstractHere we introduce the Global Spectra-Trait Initiative (GSTI), a collaborative repository of paired leaf hyperspectral and gas exchange measurements from diverse ecosystems.
Reproduction assets foundThe paper's paired leaf spectroscopy–trait database and its R processing/fitting workflow are explicitly released in a public GitHub repository, with published versions archived on ESS-DIVE.
Dataset · publicts of the GSTI will focus on expanding data coverage, incorporating data from under- represented biomes and plant functional types. 6. Data and code availability 495 The GSTI data and code are available in the public GitHub repository at https://github.com/plantphys/gsti, and published versions of GSTI are released to ESS-Dive (https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2530733, Lamour et al., 2025). 7. How to contribute to future versions of the GSTI We encourage the community to contribute new datasets to expand the scope and utility of the GSTI project. To ensure consistency and maintain data quality, contributions should adhere to the standards and guidelines outlined in this paOpen asset ↗ESS-DIVE · doi:10.15485/2530733pdf-raw-page:22 lines:1-36
Code · publicgoing refinement of spectra-trait models as new datasets are incorporated. Future developments of the GSTI will focus on expanding data coverage, incorporating data from under- represented biomes and plant functional types. 6. Data and code availability 495 The GSTI data and code are available in the public GitHub repository at https://github.com/plantphys/gsti, and published versions of GSTI are released to ESS-Dive (https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2530733, Lamour et al., 2025). 7. How to contribute to future versions of the GSTI We encourage the community to contribute new datasets to expand the scope and utility of the GSTI project. To ensure consistency and maiOpen asset ↗GitHubpdf-raw-page:22 lines:1-36
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published19 May 2025Plant phenomics (Washington, D.C.)Cited by 11 · OpenAlex ↗

PodNet: Pod real-time instance segmentation in pre-harvest soybean fields.

SoybeanField / plotRGB / grayscaleFruitSegmentationFruit / seed / panicle traits

Noninvasive analysis of pod phenotypic traits under field conditions is crucial for soybean breeding research. However, previous pod phenotyping studies focused on postharvest materials or were limited to indoor scenarios, failing to generalize to real-field environments. To address these issues, this paper employs an instance segmentation approach for the precise extraction of the pod area from multiplant RGB images in preharvest soybean fields. We first introduce a cost-effective workflow for constructing datasets of densely planted crop images with a uniform backdrop. Starting with video recording, high-quality static frames are collected by automatic selection. Then, a large vision model is explored to facilitate dense annotation and build a large-scale soybean dataset comprising 20k pod masks. Second, the pod instance segmentation model PodNet is developed based on the YOLOv8 architecture. We propose a novel hierarchical prototype aggregation strategy to fuse multiscale semantic features and a U-EMA prototype generation network to improve the model's perception performance for small objects. Comprehensive experiments suggest that lightweight PodNet achieves a superior mean average accuracy of 0.786 in the custom pod segmentation dataset. PodNet also performs competitively on in-field images without a backdrop and enables real-time inference on the edge computing platform. To the best of our knowledge, PodNet is the first pod instance segmentation model for preharvest fields. The low-cost and high-precision extraction of pods is not only a prerequisite for phenotypic analysis of the pod organs but also constitutes an important foundation in conducting cross-scale phenotyping from whole-plant to seed levels.

Why it matches plant phenotyping methods大豆莢の表現型抽出を目的に、データセット構築、インスタンスセグメンテーションモデル、実環境での性能評価を中心的に開発しているため。

abstractNoninvasive analysis of pod phenotypic traits under field conditions is crucial for soybean breeding research.
Reproduction assets foundThe authors open-source the field soybean pod instance segmentation dataset (488 images, 20k pod masks) and PodNet-related resources at their public GitHub repository, explicitly stated in the data availability statement.
Dataset · publicefficiency of manual annotation. The average pod number per image is more than 56, and the total number of pod objects is greater than 20k. Fig. 7 (c) shows that most of the pods are located in the upper center region of the image. The field soybean pod instance segmentation dataset is open sourced for the research community at https://github.com/Boatsure/PodNet . 3.2. Implementation and experiments of PodNet Considering that instance segmentation is a computationally intensive task, this study selected the lightweight architecture YOLOv8-nano (v8n) as the baseline model for the development of PodNet. Model v8n has simplified module connections and competitive perception accuracy whileOpen asset ↗Boatsure/PodNetlines:96-104
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published19 May 2025Data in briefCited by 2 · OpenAlex ↗

A comprehensive image dataset for carambola leaf and fruit disease classification and quality assessment.

Field / plotFruitLeafClassificationDisease symptoms / severity

The Carambola ( Averrhoa carambola ), also known as starfruit, is a tropical fruit with significant economic and nutritional value. The creation of a comprehensive Carambola Leaf and Fruit Dataset is highly essential for the advancement of automated disease classification and quality assessment using machine learning algorithms. The dataset is developed to support machine learning application and bridging the gap between computer vision and agricultural research to help farmers in minimizing financial losses and promote sustainable agricultural practices. The images are collected through direct field survey under diverse environmental conditions from various locations in Bangladesh between October 2024 and January 2025. The dataset comprises 2,618 original images, an equal number of processed images, and 15,000 augmented images generated from the original images. It is categorized into five distinct classes representing unique health conditions of carambola leaf and fruits including Healthy Leaves, Yellow Leaves, Insect Hole Leaves, Healthy Fruits, and Unhealthy Fruits. This dataset supports sustainable agriculture by enabling early disease identification, reduced chemical usage, and improved crop management while minimizing economic losses for farmers. It serves as a valuable resource for the future research in machine learning based plant health monitoring and quality assessment.

Why it matches plant phenotyping methodsカランボラ葉・果実の健康状態や病害を画像から分類する包括的データセットの開発であり、植物状態の画像ベース表現型評価が中心です。

abstractThe creation of a comprehensive Carambola Leaf and Fruit Dataset is highly essential for the advancement of automated disease classification and quality assessment using machine learning algorithms.
Reproduction assets foundThe paper is a Data in Brief article describing a public carambola leaf/fruit image dataset (2,618 original, processed, and 15,000 augmented images) deposited on Mendeley Data with an explicit DOI and direct URL, matching the allowed URL list.
Dataset · publicfodil Smart City, Birulia, Savar, Dhaka, Bangladesh. (Latitude: 23° 52′ 39.22" N, Longitude: 90° 19′ 12.47" E) 3. Mohamaya, Chandpur, Chittagong, Bangladesh. (Latitude: 23° 15′ 10" N, Longitude: 90° 45′ 13" E) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/f35jp46gms.1 Direct URL to data: https://data.mendeley.com/datasets/f35jp46gms/1 Related research article None 1 Value of the Data • The dataset consists of high-quality images of carambola leaves and fruits captured under different health and environmental conditions across various regions in Bangladesh. This comprehensive collection of images is a valuable resource for researchers and agronomists sOpen asset ↗Mendeley Data · 10.17632/f35jp46gms.1lines:1-54
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published19 May 2025Data in briefCited by 3 · OpenAlex ↗

PriBeL: A primary betel leaf dataset from field and controlled environment.

Field / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Essentially, visual identification of plant health is vital for research in agriculture and medicinal plants for important crops, both in terms of economics and pharmacology, such as betel leaves. The strong integration of AI-based methods in precision agriculture and herbal medicine quality control makes these systems effective only when trained on well-structured, diversified datasets.The plant betel leaf (Piper betle) is cultivated throughout the world for its medicinal, cultural, and economic importance, but improper classification and quality assessment of this plant occur because of environmental conditions and variations in handling. To solve this problem, we hereby present the Betel Leaf Dataset, which is systematically curated, consisting of 1,800 high-resolution images (1080 × 1080 pixels) exhibiting the three different conditions of betel leaves: Healthy (Fresh), Diseased, and Dried. The dataset was collected from Veer, Taluka-Purandar, Pune, India, under both natural and controlled conditions so that different appearances could be ensured. Categories include images that have been taken under varied light, backgrounds, and orientations, which comprehensively can cover all real variations in betel leaves. Hence, this systematically collected, standardized, and accessible dataset can enhance agricultural research in leaf classification studies and quality assessment techniques to facilitate better documentation and understanding of betel leaf characteristics. This dataset can be utilized in machine learning applications for plant disease detection, precision agriculture, and automated quality control systems.

Why it matches plant phenotyping methods植物の健康・病気・乾燥状態を画像化したデータセット自体が中心的な成果であり、植物状態の画像ベース表現型解析に該当する。

abstractwe hereby present the Betel Leaf Dataset, which is systematically curated, consisting of 1,800 high-resolution images (1080 × 1080 pixels) exhibiting the three different conditions of betel leaves: Healthy (Fresh), Diseased, and Dried.
Reproduction assets foundThe paper is a data descriptor for the authors' own betel leaf image dataset (1,800 images, healthy/diseased/dried, field and controlled environment), publicly deposited on Mendeley Data with an explicit direct URL and DOI.
Dataset · publicand diseased as 509. Data source location At Veer, Taluka-Purandar, District-Pune, Maharashtra, India. Latitude :18.1507784, Longitude :74.0872852 Data accessibility Repository name: Betel Leaf Dataset: A Primary Dataset From Field And Controlled Environment Data identification number: 10.17632/btdym2t6mt.1 Direct URL to data: https://data.mendeley.com/datasets/btdym2t6mt/1 Related research article None 1. Value of the Data • Betel leaves are highly grown and best known within South and Southeast Asia as piper betles due to their culinary, cultural, and medical uses. In India, the leaves are mostly grown due to warm humid conditions in states such as West Bengal, Assam, Odisha, Karnataka, TaOpen asset ↗10.17632/btdym2t6mt.1lines:1-45
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published15 May 2025PloS oneCited by 2 · OpenAlex ↗

Apple varieties, diseases, and distinguishing between fresh and rotten through deep learning approaches.

AppleFruitClassificationStress / disease detectionDisease symptoms / severity

Apples are one of the most productive fruits in the world, in addition to their nutritional and health advantages for humans. Even with the continuous development of AI in agriculture in general and apples in particular, automated systems continue to encounter challenges identifying rotten fruit and variations within the same apple category, as well as similarity in type, color, and shape of different fruit varieties. These issues, in addition to apple diseases, substantially impact the economy, productivity, and marketing quality. In this paper, we first provide a novel comprehensive collection named Apple Fruit Varieties Collection (AFVC) with 29,750 images through 85 classes. Second, we distinguish fresh and rotten apples with Apple Fruit Quality Categorization (AFQC), which has 2,320 photos. Third, an Apple Diseases Extensive Collection (ADEC), comprised of 2,976 images with seven classes, was offered. Fourth, following the state of the art, we develop an Optimized Apple Orchard Model (OAOM) with a new loss function named measured focal cross-entropy (MFCE), which assists in improving the proposed model's efficiency. The proposed OAOM gives the highest performance for apple varieties identification with AFVC; accuracy was 93.85%. For the apples rotten recognition with AFQC, accuracy was 98.28%. For the identification of the diseases via ADEC, it was 99.66%. OAOM works with high efficiency and outperforms the baselines. The suggested technique boosts apple system automation with numerous duties and outstanding effectiveness. This research benefits the growth of apple's robotic vision, development policies, automatic sorting systems, and decision-making enhancement.

Why it matches plant phenotyping methodsリンゴ画像から腐敗状態や病害を推定するデータセットと深層学習モデルを開発・評価しており、植物状態の画像ベース推定が中心である。

abstractwe first provide a novel comprehensive collection named Apple Fruit Varieties Collection (AFVC) with 29,750 images through 85 classes.
Reproduction assets foundThe paper's three apple image datasets (AFVC, ADEC, AFQC) are explicitly released with free public access via the authors' GitHub repositories, and the Data Availability statement confirms all data is available at these URLs. These are paper-specific image datasets used directly for the paper's apple variety, disease,,
Dataset · public7) 2,682 294 2,976 Fig 5 The Apple Fruit Varieties Collection (AFVC) distributions through 85 classes. Fig 6 The Apple Fruit Varieties Collection (AFVC) measurement was split through 85 classes; the overall training was 26,775, and the testing was 2,975 samples. The second collection, Apple Fruit Quality Categorization (AFQC) [ https://github.com/mustafa20999/AFQC ], was collected from the orchard ( Table 1 ). The study area was Beijing City, Huairou District, Beijing Shengshiguowang, with a mean temperature of 76°C − 19°C and an average monthly rainfall of 51.2 mm. Data was collected at two different periods between October 1st, 2023, and October 10th, 2023: in the morning, when shootinOpen asset ↗mustafa20999/AFQClines:66-100
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published14 May 2025Data in briefCited by 2 · OpenAlex ↗

Comprehensive dataset on ripening stages of strawberries and avocados: From unripe to rotten.

AvocadoStrawberryFruitClassificationObject detectionGrowth / development / phenology

This paper presents a novel and innovative 14,630 fruit images dataset, consisting of 1333 original images and the remaining augmented images for strawberry and avocado fruits. The dataset records the growth of strawberries and avocados in four different stages: unripe, partially ripe, ripe, and rotten. Though the fruit ripening process is commonly known, a lack of systematic datasets to show the fruit changing from an unripe state to a rotting state was prevalent for the two fruits in question. Over two months, the dataset was collected through rigorous tracking to effectively provide a measure of each of the fruits' conditions. The fruits were obtained from Mahabaleshwar farms in Maharashtra, India, as well as from local markets in Maharashtra and Pune. The fruits were monitored continuously from the time of harvesting, and all observed changes were carefully recorded. The uniqueness of this dataset is that it covers both strawberries and avocados, which have different patterns of ripening and are highly commercially valuable. The images were annotated using the online annotation tool - makesense.ai, with a total of 1499 bounding boxes for each fruit. By encompassing these two diverse fruit types, the dataset provides a valuable resource for researchers, agriculturalists, and food scientists to investigate and compare the ripening behaviours of different fruit species.

Why it matches plant phenotyping methodsイチゴとアボカドの果実画像を用いて、未熟から腐敗までの可視的な成熟・状態を体系的に記録し、注釈付きデータセットとして提供しているため、植物器官の状態を対象とする画像ベースのフェノタイピングデータセットに該当する。

abstractThis paper presents a novel and innovative 14,630 fruit images dataset, consisting of 1333 original images and the remaining augmented images for strawberry and avocado fruits.
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository containing the authors' own fruit image dataset (14,630 strawberry/avocado images with YOLO bounding-box annotations across ripening stages), which directly constitutes the paper's phenotyping measurements. No analysis code or trained模型s是
Dataset · publicset up using a white background for enabling consistent and uniform image acquisition.. Data source location Dataset was collected from (i) Mahabaleshwar, Maharashtra, India; and (ii) Pune, Maharashtra, India. Data accessibility Repository name: mendeley.com Data identification number: 10.17632/zysvgmxcyz.1 Direct URL to data: https://data.mendeley.com/datasets/zysvgmxcyz/1 Related research article 1. Value of the Data • This dataset is a useful resource for machine learning solutions in fruit maturity detection and can contribute to the design of automated sorting and classification systems by ripeness stages. • Food processing companies and agricultural scientists may utilize this data to Open asset ↗10.17632/zysvgmxcyz.1lines:1-51
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 May 2025Data in briefCited by 2 · OpenAlex ↗

Near-Infrared Spectroscopy and Wet Chemistry Dataset for Forage Nutritional Quality Assessment in Urochloa humidicola .

Field / plotRaman / spectroscopyWhole plant / canopy / plot / fieldPhysiological trait estimation

Assessing the nutritional quality traits of pastures is crucial for germplasm and breeding evaluations, enabling the selection of high-quality forages to enhance livestock productivity. However, traditional laboratory analytical methods are logistically demanding and costly, particularly in large-scale trials, underscoring the need for rapid, precise, and high-throughput evaluation methods. Near-Infrared Spectroscopy (NIRS) optimizes the estimation of forage nutritional quality parameters by developing chemometric models that predict these parameters with high accuracy and precision, based on the association between NIRS data and wet chemistry analyses. This dataset, collected over ten years by the Tropical Forages Program at the International Center for Tropical Agriculture (CIAT) in Colombia, comprises 1112 samples. It includes 995 measurements of Neutral Detergent Fiber (NDF), 996 of Acid Detergent Fiber (ADF), 995 of In Vitro Dry Matter (IVDMD), and 469 of Crude Protein (CP), all obtained through wet chemistry methodologies. Additionally, the 1112 samples contain absorbance data spanning 400 to 2498 nanometers (nm) in 2 nm intervals, generating 1050 spectral data points per sample. Finally, this dataset is a valuable resource for predicting forage nutritional quality beyond conventional parameters, incorporating plant reflectance attributes to enhance selection strategies for optimized forage selection.

Why it matches plant phenotyping methods牧草の栄養品質形質をNIRSスペクトルから推定するための大規模データセットであり、湿式化学値との対応付けとケモメトリックモデル構築が中心的な方法的貢献である。

abstractNear-Infrared Spectroscopy (NIRS) optimizes the estimation of forage nutritional quality parameters by developing chemometric models that predict these parameters with high accuracy and precision, based on the association between NIRS data and wet chemistry analyses.
Reproduction assets foundThe paper is a data descriptor for a paper-specific public dataset: 1112 Urochloa humidicola samples with wet-chemistry traits (NDF, ADF, IVDMD, CP) and 1050-point NIR absorbance spectra (400–2498 nm), deposited in the Harvard Dataverse (DOI 10.7910/DVN/XPNIQY). The deposit is explicitly public and actionable; however,
Dataset · publicData accessibility Repository name: Harvard database Data identification number: 10.7910/DVN/XPNIQYHarvard database · 10.7910/DVN/XPNIQYlines:1-49
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published7 May 2025Data in briefCited by 2 · OpenAlex ↗

A comprehensive image dataset of plum leaf and fruit for disease classification.

PlumFruitLeafClassificationDisease symptoms / severity

Plums, commonly known as Indian jujube, are economically important, valued for nutritional benefits and consumed by people from all over the world. The development of a comprehensive Plum leaf and fruit dataset is highly essential for advancing agricultural research and enabling effective disease management systems using machine learning techniques. This dataset serves as a foundational resource for machine learning based classification and bridges the gap between agricultural research and computer vision to support automated disease detection and fruit quality assessment. Researchers will be able to utilize this dataset to implement early disease detection which leads to improve crop management and supply quality and reduce the usage of chemicals. Proper utilization of this dataset can help farmers to reduce financial losses and encourage sustainable farming practices. The dataset was collected between December 2024 and February 2025 under various environmental conditions. It consists of 3,554 original images, an equal number of processed images and 18,000 augmented images generated from the original dataset. The dataset is categorized into six distinct classes: Shot Hole, Bacterial Spot, Wilted Leaf, Healthy Leaf, Unhealthy Plum, and Healthy Plum. This dataset contributes significantly to advance deep learning in agriculture enabling early disease detection and fruit quality monitoring.

Why it matches plant phenotyping methods植物の葉・果実画像から病害状態を分類するためのデータセット構築が研究の中心であり、植物病害フェノタイピング用の再利用可能な資源に該当する。

abstractThis dataset contributes significantly to advance deep learning in agriculture enabling early disease detection and fruit quality monitoring.
Reproduction assets foundThe paper is a Data in Brief describing a public Mendeley Data repository containing the authors' own plum leaf/fruit image dataset (3,554 original, processed, and 18,000 augmented images) used for plant disease classification. This is a paper-specific, publicly available image dataset with an explicit URL and DOI.
Dataset · publicil Smart City, Birulia, Savar, Dhaka, Bangladesh. (Latitude: 23° 52′ 39.22" N, Longitude: 90° 19′ 12.47" E) 3. Sharankhola, Bagerhat, Khulna, Bangladesh (Latitude: 22° 13′ 26.0" N, Longitude: 89° 48′ 20.0" E). Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/w7sdx55m7z.1 Direct URL to data: https://data.mendeley.com/datasets/w7sdx55m7z/1 Related research article none 1 Value of the Data •Open asset ↗Mendeley Data · 10.17632/w7sdx55m7z.1lines:1-52
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published7 May 2025Data in briefCited by 0 · OpenAlex ↗

Detecting olive quick decline syndrome: A satellite-based dataset for a case study in Apulia Region.

OliveAerial / UAVField / plotMultispectral / hyperspectralWhole plant / canopy / plot / fieldSegmentationStress / disease detectionDisease symptoms / severity

The bacterium Xylella fastidiosa (Xf) is a plant pathogen first identified in Europe in 2013, specifically in olive groves in the Apulia region (south-eastern Italy). It is now spreading across the Mediterranean basin and poses a serious threat to the local economy by causing branch desiccation and the rapid death of olive trees, a condition known as olive quick decline syndrome (OQDS). Several studies have investigated the potential of remote sensing (RS) technology to monitor OQDS over time and space; however, accurate and reliable data on OQDS occurrence remain scarce. To enhance the distribution data of Xf-infected trees in the Apulia region, we investigated an infection hotspot of 25 km² area in the province of Brindisi, where records of infections were documented in 2019 and 2020. Three very high resolution, commercial WorldView-2 images were acquired and segmented, resulting in a dataset of 76637 olive trees. Through visual interpretation, 2340 trees were identified most likely as either infected or removed due to OQDS. This dataset provides a valuable resource for developing or validating RS techniques for early detection of OQDS. Furthermore, it could support studies aimed to evaluate spectral bands or indices most correlated with infection presence. Finally, the dataset can be integrated with other Xf-infection presence data to support species distribution model studies.

Why it matches plant phenotyping methods衛星画像のセグメンテーションと感染・枯死オリーブ樹のラベル化による、植物病害状態の検出・検証用データセットが研究の中心である。

abstractThree very high resolution, commercial WorldView-2 images were acquired and segmented, resulting in a dataset of 76637 olive trees.
Reproduction assets foundThe paper is a Data in Brief article describing a public Figshare dataset (OQDS-Insight) containing WorldView-2 satellite raster imagery (RGB and NDVI GeoTIFFs) and a shapefile of 76,637 olive tree points with OQDS infection labels — directly the paper's phenotyping measurements.
Dataset · publicsouth-eastern Italy. The extent (EPSG:32633) is from 706164.541 N to 713395.999 N, and from 4508710.411 E to 4513574.414 E. Data are stored at the Council for Agricultural Research and Economics, Research Centre for Agriculture and Environment, Italy. Data accessibility Repository name: OQDS-Insight Data identification number: https://doi.org/10.6084/m9.figshare.28191245.v4 Direct URL to data: https://doi.org/10.6084/m9.figshare.28191245.v4 Related research article None. Open in a new tab 1. Value of the Data • The dataset provides a detailed record of OQDS olive groves within an infection hotspot in the province of Brindisi, Apulia region (south-eastern Italy) ( Fig. 1 ). • It can support rOpen asset ↗figshare · 10.6084/m9.figshare.28191245.v4lines:95-140
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 May 2025Food science & nutritionCited by 4 · OpenAlex ↗

A Lightweight Framework for Protected Vegetable Disease Detection in Complex Scenes.

GreenhouseObject detectionStress / disease detectionDisease symptoms / severity

The rapid development of computer vision technology has provided new technical support for smart agriculture. Vegetable diseases represent a significant threat to agricultural production, with severity that cannot be ignored. However, through scientifically effective prevention and control measures, these negative impacts can be significantly mitigated. Intelligent disease detection systems, as advanced methods replacing traditional manual inspection, have become important means for developing smart agriculture and improving the efficiency of vegetable production management. Nevertheless, traditional manual detection is not only time-consuming and labor-intensive but also faces accuracy limitations, while existing computer vision detection methods still encounter a series of challenges when confronting complex backgrounds, diverse disease manifestations, and varying degrees of occlusion in real cultivation environments, including insufficient anti-interference capabilities, limited detection precision, and suboptimal real-time performance. This research addresses the practical challenges of limited data acquisition and sample scarcity for protected vegetable diseases by proposing an innovative strategy that implements differentiated data augmentation technique combinations for different categories of samples, significantly enhancing the model's resistance to environmental interference. Based on the integrated concepts of machine vision and deep learning, we developed a lightweight vegetable disease detection network named VegetableDet. This network innovatively combines Deformable Attention Transformer (DAT) with YOLOv8n backbone architecture, enhancing perception capabilities for long-range feature dependencies. Simultaneously, a Channel-Spatial Adaptive Attention Mechanism (CSAAM) is integrated into the Neck network, achieving precise localization and enhancement of key features. To address the issue of low model convergence efficiency, we further designed a hierarchical progressive transfer learning training strategy, effectively accelerating the model adaptation process and improving detection accuracy. Experimental evaluation demonstrates that on our custom comprehensive protected vegetable disease dataset, the VegetableDet model exhibits excellent performance in detecting 30 diseases and healthy samples across 5 vegetable types, with precision (P), recall (R), and average precision (AP) all exceeding 90%, and an overall mean Average Precision (mAP) reaching 94.31%. The model demonstrates powerful adaptability under complex environmental conditions, providing reliable technical support for real-time monitoring and precise prevention and control of protected vegetable diseases, with broad application prospects.

Why it matches plant phenotyping methods植物病害の症状を画像から検出・分類する軽量深層学習手法とデータセットを開発し、複雑環境で性能評価しており、植物状態の取得・推定が中心である。

abstractwe developed a lightweight vegetable disease detection network named VegetableDet.
Reproduction assets foundThe paper's Data Availability Statement points to a public GitHub repository containing part of the self-collected protected vegetable disease detection dataset (and code), with the complete dataset/code available on request from the corresponding author.
Dataset · publicata Availability Statement The data utilized in this paper is obtained through self‐gathering and is made publicly available (a part of it) to make the study reproducible. The datasets generated and analyzed during the current study are partly available in the github repository, accessible via the following persistent web link: https://github.com/tyuiouio/plant‐disease‐detection‐in‐real‐field . If you want to request the complete dataset and code, please email the corresponding author. References Attri, I. , Awasthi L. K., and Sharma T. P.. 2025. “EQID: Entangled Quantum Image Descriptor an Approach for Early Plant Disease Detection.” Crop Protection 188: 107005. Bao, W. , Zhu Z., Hu G., ZhoOpen asset ↗tyuiouio/plant‐disease‐detection‐in‐real‐fieldlines:559-618
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published1 May 2025Applications in plant sciencesCited by 0 · OpenAlex ↗

Improving computer vision for plant pathology through advanced training techniques.

Cocoa / cacaoWhole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

Premise This study investigates advanced training techniques to improve the performance of convolutional neural networks for disease detection in cocoa, Theobroma cacao . Methods Despite recent stagnation in accuracy improvements in computer vision for image classification, our research demonstrates significant advancements in performance through semi-supervised learning, specialised loss functions, and the inclusion of a non-cocoa class. Results Semi-supervised learning reduced overfitting and enhanced generalisability, particularly for subtle symptoms. The non-cocoa class exposed models to a broad range of relevant features, significantly improving model robustness and performance in difficult cases. Grad-CAM for qualitative assessment provided valuable insights into model behaviour, highlighting cases of overfitting missed by summary statistics. We also describe dynamic focal loss, a novel loss function that uses an empirical measure of difficulty to weight each image. Our results suggest that while PhytNet shows promise in terms of computational efficiency and superior handling of difficult images, ResNet18 with semi-supervised learning and dynamic focal loss emerged as the strongest contender for real-world deployment. Discussion This research underscores the potential of semi-supervised learning and advanced loss functions in enhancing the applicability of deep learning models in agricultural disease management. It also presents a new high-quality benchmark dataset of 7220 images of diseased and healthy cocoa trees, offering a much greater and more realistic challenge than the Plan Village dataset.

Why it matches plant phenotyping methodsカカオ葉・樹体の病徴画像から植物の病害状態を推定する深層学習手法を開発・比較し、性能評価とベンチマークデータセット構築を行っており、フェノタイピング手法が中心である。

abstractThis study investigates advanced training techniques to improve the performance of convolutional neural networks for disease detection in cocoa, Theobroma cacao .
Reproduction assets foundThe paper's data availability statement explicitly provides the paper-specific cocoa image dataset and the FAIGB dataset on OSF, plus authors' analysis code on GitHub, all with public URLs.
Dataset · publicThe cocoa image data is available at https://osf.io/2fw6gOpen asset ↗osf · 2fw6glines:753-923
Dataset · publicthe FAIGB web‐scraped dataset of crop disease images is available at https://osf.io/nuafhOpen asset ↗osf · nuafhlines:753-923
Code · publicAll code necessary to reproduce these results is available on GitHub ( https://github.com/jrsykes/CocoaReader/tree/main/CocoaNet/PhytNet_Cocoa )Open asset ↗github · jrsykes/CocoaReaderlines:753-923
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published1 May 2025EcologyCited by 2 · OpenAlex ↗

TropiRoot 1.0: Database of tropical root characteristics across environments.

RootVisualization / data managementBiomass / plant weightRoot system architecture

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

Why it matches plant phenotyping methods熱帯植物の根形態・構造・生理などの表現型特性を標準化して収録した再利用可能なデータベースであり、データセット構築と品質管理が中心です。

abstractHere, we provide a database of tropical root characteristics.
Reproduction assets foundThe paper's core asset is the TropiRoot 1.0 root trait database itself, publicly deposited in ESS-DIVE (DOI 10.15485/2507279) and also provided as Supporting Information (Data S1). This is a paper-specific public phenotype/trait dataset directly reproducing the paper's measurements.
Dataset · publich, et al. 2025. “ TropiRoot 1.0: Database of Tropical Root Characteristics across Environments.” Ecology 106(5): e70074. 10.1002/ecy.70074 Handling Editor: Simona Picardi DATA AVAILABILITY STATEMENT The dataset is available as Supporting Information to this Ecology data paper and is also accessible in the ESS‐DIVE repository at https://doi.org/10.15485/2507279. Associated Data Supplementary Materials Data S1. Data Availability Statement The dataset is available as Supporting Information to this Ecology data paper and is also accessible in the ESS‐DIVE repository at https://doi.org/10.15485/2507279.Open asset ↗10.15485/2507279html-lines:63-80
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published29 Apr 2025Data in briefCited by 2 · OpenAlex ↗

Comprehensive smartphone image dataset for Aegle Marmelos, Hog plum, and lemon plant leaf disease and freshness assessment.

CitrusField / plotLeafClassificationDisease symptoms / severity

Fruits, which are packed with nutrients, vitamins, and antioxidants, have been known for their numerous health benefits and curative powers, and are utilized in conventional medicine. Aegle Marmelos, Lemon, and Hog Plum are tangy fruits widely recognized in Asian countries for containing a plentiful supply of bioactive substances. They are also highly valuable in boosting metabolism, possessing tremendous therapeutic properties, and holding financial significance. The leaves of these fruit trees are as essential as their fruits, as they contain versatile medicinal and dietary benefits of immense value. However, these leaves are often affected by various fungal and other diseases, which reduce the ability for healthy growth and productivity of both fruits and leaves. Plants infected with various leaf diseases can produce fewer fruits, which are also of lower quality due to failure to reach maturity and lack of sufficient nutritional value. For these reasons, there is a risk of an outbreak in orchards, which can lead to significant financial losses for both producers and the agricultural sector. This signifies that the early identification of leaf diseases and the management of orchards are essential to minimize the impact of leaf diseases and mitigate these issues, ensuring the healthy production of valuable medicinal fruits. In this paper, various infected leaf images are collected from different regions of Rangpur, providing a comprehensive dataset comprising 3941 images. The dataset includes images of three different plant leaves, where 1513 images of Aegle Marmelos, 1232 images of Lemon, and 1196 of Hog plum, where each of the categories encompasses several classes of common leaf diseases. Through this dataset, an early and accurate digital detection system can be employed, allowing producers to clearly identify diseases instead of relying on traditional methods. The precise and timely identification of leaf diseases enables the control of these diseases by taking necessary actions, ensuring the sustainability of plants, and promoting the healthy growth of these invaluable medicinal fruits.

Why it matches plant phenotyping methods植物葉の病害状態を画像で評価する大規模データセットの構築が中心であり、病害表現型の画像ベース解析に該当する。

abstractIn this paper, various infected leaf images are collected from different regions of Rangpur, providing a comprehensive dataset comprising 3941 images.
Reproduction assets foundThe paper's own smartphone leaf-image dataset (3,941 raw + 12,295 augmented images of Aegle Marmelos, Hog plum, and lemon leaves) is publicly deposited on Mendeley Data with an explicit direct URL and DOI, matching the allowed URL list. No separate analysis code or trained model checkpoint is stated as available.
Dataset · publicden in Rangpur district (latitude: 25° 34′ 30.6942″, longitude: 89° 16′ 22.2672″), 2. Chotali Kosba Para village fruits garden in Rangpur district (latitude: 25° 34′ 30.6942″, longitude: 89° 16′ 22.2672″) Data accessibility Repository name: Mendeley Data Data identification number: DOI: 10.17632/54r883j5zr.1 Direct URL to data: https://data.mendeley.com/datasets/54r883j5zr/1 1 Value of the Data •Open asset ↗Mendeley Data · 10.17632/54r883j5zr.1lines:1-45
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published28 Apr 2025Data in briefCited by 2 · OpenAlex ↗

Cauliflower leaf diseases: A computer vision dataset for smart agriculture.

Brassica vegetablesLeafClassificationStress / disease detectionDisease symptoms / severity

Cauliflower is among the more well-known vegetables there are. Consumed all around the globe due to it being rich in nutrients such as vitamins, antioxidants, and for being high in fibre. These are nutritional qualities that help with digestion, immune-system, and minimizing inflammation. It is a common issue among farmers to have to deal with various diseases in cauliflower leaves that are difficult to diagnose in their early stages. These diseases have a tendency to propagate in a really swift pace throughout entire fields worth of crops. This in-turn causes heavy losses in the harvest, and makes it much more tedious and resource-intensive to protect the crops. As a result, farmers get more likely to use high amounts of pesticides and harmful chemicals to streamline the process of getting a more reliable yield on their crops. This is not only costly, but it is also harmful both to the quality of crops and to the well-being of the environment. In this publication, we are introducing a dataset containing a considerable number of images of cauliflower leaves. This is intended to drive development on this topic at a faster pace than it is now, and to help enhance disease monitoring, diagnosis, and precautionary techniques. We collected our dataset images between November 2024 and January 2025. In this dataset, cauliflower leaves were categorized into three classes: Healthy, Insect Holes, and Black Rot, each reflecting a specific condition that impacts plant health at different stages. This dataset consists of 2,661 images. The pictures were captured at different locations in Bangladesh, under different weather conditions, dates, temperatures, and with different devices. To enhance the data quality, we used several steps to process the dataset, making sure it would reflect real-world conditions and be ready for training. The images were resized to a standard size of 3000 × 3000 pixels, brightness was adjusted to make the images more easily discernible, and we removed duplicates and poor-quality images. These actions helped ensure the dataset was in the best possible shape for effective model training. This dataset will be highly effective for agricultural research, precision agriculture, and effective management of diseases. It should help develop highly accurate machine learning models for early detection of Cauliflower leaf diseases. The dataset is employed to train deep learning models to support automated monitoring and smart decision-making in precision agriculture. This data set also has immense potential for real-time and practical use. It can be utilized to develop applications like mobile apps or automated systems where farmers can easily identify diseases at early stages and take immediate action, without the requirement of expert on-site knowledge. This data set can also be utilized with smart farming equipment like drones and sensors to track big fields in real time.

Why it matches plant phenotyping methodsカリフラワー葉の健康状態・病害状態を画像で分類するデータセット自体が研究の中心であり、植物病害表現型の取得・解析基盤に該当します。

abstractIn this publication, we are introducing a dataset containing a considerable number of images of cauliflower leaves.
Reproduction assets foundThe paper's core asset is its own cauliflower leaf disease image dataset (2,661 images, three classes), publicly deposited on Mendeley Data with DOI 10.17632/x995snz7p3.1 and a direct URL matching an allowed URL.
Dataset · publiced from the following geographic locations: 1. Zailla, Singair, Manikganj Latitude : 23°47′46.11"N Longitude : 90°13′15.73"E 2. Dattapara, Ashulia, Savar, Dhaka Latitude : 23°52′26.3"N Longitude : 90°19′06.3"E Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/x995snz7p3.1 Direct URL to data: https://data.mendeley.com/datasets/x995snz7p3/1 The dataset is publicly available and can be accessed via the provided Mendeley Data repository link. Related research article None 1. Value of the Data • This dataset holds high-resolution images of diseased cauliflower leaves infected with multiple diseases, which provide a wealth of material for the development and vaOpen asset ↗Mendeley Data · 10.17632/x995snz7p3.1lines:35-107
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published28 Apr 2025Data in briefCited by 0 · OpenAlex ↗

Updating high-resolution image dataset for the automatic classification of phenological stage and identification of racemes in Urochloa spp. hybrids with expanded images and annotations.

Field / plotRGB / grayscalePanicle / ear / spikeClassificationObject detectionGrowth / development / phenology

This dataset is an expanded version of a previously published collection of high-resolution RGB images of Urochloa spp. genotypes, initially designed to facilitate automated classification of phenological stages and raceme identification in forage breeding trials. The original dataset included 2400 images of 200 genotypes captured under controlled conditions, supporting the development of computer vision models for High-Throughput Phenotyping (HTP). In this updated release, 139 additional images and 24,983 new annotations have been added, bringing the dataset to a total of 2539 images and 47,323 raceme annotations. This version introduces increased diversity in image-capture conditions, with data collected from two geographic locations (Palmira, Colombia, and Ocozocoautla de Espinosa, Mexico) and a range of image-capture devices, including smartphones (e.g. Realme C53 and Oppo Reno 11), a Nikon D5600 camera, and a Phantom 4 Pro V2 drone. Images now vary in perspective (nadir, high-angle, and frontal) and capture distance (1-3 meters), enhancing the dataset applicability for robust Deep Learning (DL) models. Compared to the original dataset, raceme density per plant has nearly doubled in some samples, offering higher raceme overlap for advanced instance segmentation tasks. This expanded dataset supports deeper exploration of phenotypic variation in Urochloa spp. and offers greater potential for developing adaptable models in crop phenotyping.

Why it matches plant phenotyping methods植物の生育ステージ分類と総状花序の同定を目的とする画像データセットで、注釈付き画像の拡張、撮影条件の多様化、インスタンスセグメンテーション用途が中心であり、再利用可能な表現型解析基盤に該当する。

abstractThis dataset is an expanded version of a previously published collection of high-resolution RGB images of Urochloa spp. genotypes, initially designed to facilitate automated classification of phenological stages and raceme identification in forage breeding trials.
Reproduction assets foundThe paper is a Data in Brief describing a public Harvard Dataverse deposit of the paper's own Urochloa spp. hybrid RGB images and COCO raceme annotations, with a direct DOI URL listed in allowed_urls.
Dataset · publicupo Papalotla City 1: Palmira, Valle del Cauca. City 2: Ocozocoautla de Espinosa, Chiapas. Country 1: Colombia. Country 2: Mexico. Geolocalization 1: 3°29’N, 76°21’W Geolocalization 2: 16°45′N 93°28′W Data accessibility Repository name: Harvard Dataverse Data identification number: doi.org/10.7910/DVN/X4LM19 Direct URL to data: https://doi.org/10.7910/DVN/X4LM19Open asset ↗Harvard Dataverse · 10.7910/DVN/X4LM19lines:1-51
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published22 Apr 2025Data in briefCited by 6 · OpenAlex ↗

A comprehensive image dataset for accurate diagnosis of betel leaf diseases using artificial intelligence in plant pathology.

Field / plotLeafStress / disease detectionDisease symptoms / severity

In South Asian countries, agriculture is a crucial employment field, and a remarkable number of people depend on it for their livelihood. Crop diseases are a significant threat to sustainable development in the agriculture field. Automated efficient crop disease diagnosis techniques developed with comprehensive field image datasets can play a vital role in preventing diseases at an early stage. Betel leaf is widely consumed in South Asian countries for its nutritional benefits, but to the best of our knowledge, no extensive dataset of betel leaf is available that can play a crucial role in developing accurate disease diagnosis tools. Farmers face a significant economic loss due to betel leaf diseases, and due to the lack of efficient diagnosis tools, the farming of betel leaf has become very difficult day by day. Our motive is to develop a reliable and versatile image dataset of field images that will assist artificial intelligence-based pathology research on betel leaf diseases. This dataset contains healthy leaf images and two common disease images of betel leaf such as leaf rot and leaf spot [1]. Initially, 2,037 betel leaf images were captured in a natural daylight environment from several betel cultivation fields in Bangladesh. Afterward, 10,185 images were generated using image augmentation strategies including flipping, brightness factor, contrast factor, and rotation. This dataset is well-compatible with machine learning and deep learning-based pathology research, as it contains enough image samples for model training, validation, and testing. Moreover, a comparison study is conducted that ensures this dataset fulfills the gap of a reliable and extensive dataset of betel leaf. This comprehensive dataset serves as a crucial resource for researchers in developing efficient computational models for accurate betel leaf disease diagnosis.

Why it matches plant phenotyping methodsベテル葉の病害状態を画像で収録したデータセットの開発が中心であり、植物病害の画像ベース表現型評価に直接利用できる。

abstractOur motive is to develop a reliable and versatile image dataset of field images that will assist artificial intelligence-based pathology research on betel leaf diseases.
Reproduction assets foundThe paper's core asset is its own betel leaf disease image dataset (2,037 original + 10,185 augmented images), publicly deposited on Mendeley Data with an explicit direct URL and DOI.
Dataset · public1. Charshihari, Ishwarganj, Mymensingh, Bangladesh. 2. Chorhossianpur, Ishwarganj, Mymensingh, Bangladesh. 3. Uchakhila, Ishwarganj, Mymensingh, Bangladesh. 4. Lakshmiganj, Ishwarganj, Mymensingh, Bangladesh. Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/vpzkntzjty.1 Direct URL to data: https://data.mendeley.com/datasets/vpzkntzjty/1 The dataset can be accessed directly using the mentioned URL, a zip file of 1.28 GB will be downloaded. Researchers or interested individuals can use the images of the dataset by extracting the zip file. Related research article None 1. Value of the Data • This unique dataset of betel leaf images is crucial for the crop-bOpen asset ↗Mendeley Data · 10.17632/vpzkntzjty.1lines:1-49
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published8 Apr 2025Data in briefCited by 6 · OpenAlex ↗

Irish potato imagery dataset for detection of early and late blight diseases.

PotatoField / plotLeafClassificationDisease symptoms / severity

This dataset comprises of 58,709 annotated images of irish potato leaves, categorized into three classes (healthy, early blight and late blight). The data was collected over six months from smallholder farms in Southern Highlands Tanzania, using Samsung Galaxy A03 smartphones with 8-megapixel camera. Researchers, farmers and agricultural extension officers were trained to capture images under diverse conditions, including varying lighting, angles and backgrounds to ensure the dataset is diverse and representative. Plant pathologists were used to validate the images to ensure and enhance the reliability of the labels. Pre-processing steps such as duplicate removal, filtering of irrelevant images, annotation and metadata integration were applied resulting in a high-quality dataset. The dataset is organized into three folders (healthy, early blight and late blight) and is freely available on the Zenodo repository to promote accessibility for researchers working in the field of plant diseases. This dataset holds significant potential for reuse in training machine learning models for crop disease detection, transfer learning and data augmentation studies. By enabling early detection and classification of potato diseases, the dataset supports the development of innovative agricultural tools aimed at reducing crop losses and enhancing food security in Sub-Saharan Africa. Its robust design and regional specificity make it a valuable resource for advancing research and innovation in sustainable farming practices.

Why it matches plant phenotyping methodsジャガイモ葉の画像から健全・初期疫病・後期疫病という植物病害状態を判定する、注釈付き大規模データセットであり、再利用可能なフェノタイピング資源として中心的です。

abstractThis dataset comprises of 58,709 annotated images of irish potato leaves, categorized into three classes (healthy, early blight and late blight).
Reproduction assets foundThe paper is a data descriptor for the authors' own Irish potato leaf imagery dataset (58,709 annotated images for healthy/early blight/late blight classification), publicly deposited on Zenodo with an explicit DOI and direct URL matching an allowed URL. This is a paper-specific public plant image/phenotyping asset.
Dataset · publiccollected from farms located in Southern Highlands of Tanzania, specifically in Mbeya (8.9090° S, 33.4589° E), Iringa (7.7673° S, 35.6900° E), Njombe (9.3333° S, 34.7667° E) and Songwe (9.1333° S, 32.9333° E) regions. Data accessibility Repository name: ZENODOData identification number: 10.5281/zenodo.8286529Direct URL to data: https://zenodo.org/records/8286529 1. Value of the Data • This dataset serves as a valuable resource as it addresses critical data gaps in the field of Artificial Intelligence in Agriculture by providing region-specific dataset with over 58,000 annotated images captured under diverse environment in the real-world smallholder farming conditions. • This dataset caOpen asset ↗Zenodo · 10.5281/zenodo.8286529html-lines:1-30
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Apr 2025Data in briefCited by 2 · OpenAlex ↗

A comprehensive Malabar Spinach dataset for diseases classification.

LeafClassificationStress / disease detectionDisease symptoms / severity

This study focuses on the urgent need to increase detection of diseases in Malabar Spinach, a valuable leaf vegetable crop which is at risk from several disease types including Anthracous leaf spot and Straw mite infestation. There is still a lack of research focused on Malabar spinach, although advances in machine vision have considerably increased the detection of largescale crop diseases. By developing and evaluating machine vision algorithms specifically designed for accurate detection of diseases in Malabar spinach, this research aims to fill this gap. To achieve this, a comprehensive dataset comprising images of both healthy and diseased Malabar Spinach plants is utilized for training, testing, and validation purposes. This study seeks to develop reliable disease detection models through the examination of different image processing techniques and deep learning algorithms such as ResNet50. In particular, the performance of these models is rigorously evaluated on the basis of a set of standardized evaluation metrics which aim to achieve an overall test accuracy of 94%. The results of this research will have a major impact on the cultivation of Malabar spinach in terms of precision farming techniques and effective crop management practices. This study will contribute to the wider objectives of agricultural sustainability and food security, through increasing crop productivity and reducing yield losses. In the end, it is intended to strengthen the resilience of farming communities dependent on Malabar Spinach crops by providing farmers and experts with efficient tools for detecting diseases.

Why it matches plant phenotyping methodsマラバルホウレンソウの健全・罹病状態を画像から分類するデータセットと画像処理・深層学習手法の開発および評価が研究の中心であり、植物病害状態のフェノタイピングに該当する。

abstractBy developing and evaluating machine vision algorithms specifically designed for accurate detection of diseases in Malabar spinach
Reproduction assets foundThe paper's own Malabar Spinach disease image dataset (603 original + 5868 augmented images) is publicly deposited on Mendeley Data with DOI 10.17632/n56pn9fncw.2 and a direct URL, making it a paper-specific, publicly accessible phenotyping image asset.
Dataset · publica for training and testing. Data source location Location: Local Agriculture field. Zone: Birulia, Ashulia, Savar, Dhaka. Country: Bangladesh Data accessibility Repository name: Malabar Spinach dataset for diseases classification using deep learning approach. Data identification number: 10.17632/n56pn9fncw.2 Direct URL to data: https://data.mendeley.com/datasets/n56pn9fncw/2 1. Value of the Data •Open asset ↗10.17632/n56pn9fncw.2lines:1-44
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 Mar 2025Data in briefCited by 11 · OpenAlex ↗

Tomato leaf dataset: A dataset for multiclass disease detection and classification.

TomatoLeafClassificationDisease symptoms / severity

Agriculture is a cornerstone of Bangladesh's economy, with tomatoes being one of the most widely cultivated vegetables, producing approximately 368,000 tons annually. However, tomato plants are vulnerable to various diseases and pest infestations that can significantly reduce crop yield, posing a threat to farmers' livelihoods. Early detection of these diseases, often visible through symptoms on the leaves, is critical for effective management. In this work, we present a dataset of 731 high-resolution images of tomato leaves affected by six common diseases, along with healthy samples, aimed at facilitating automated disease diagnosis using computer vision. The dataset is categorized into disease types such as Early Blight, Black Spot, Late Blight, Leaf Mold, Bacterial Spot, and Target Spot. This structured dataset offers a valuable resource for researchers developing machine learning models for disease classification and early detection. By making the dataset publicly available, we aim to accelerate research in precision agriculture and empower the development of AI-driven tools that can enhance tomato disease management, ultimately improving crop yields and supporting sustainable farming practices.

Why it matches plant phenotyping methodsトマト葉の病徴画像を収録した公開データセットであり、植物の病害状態を画像から分類するフェノタイピング用資源が中心です。

abstractIn this work, we present a dataset of 731 high-resolution images of tomato leaves affected by six common diseases, along with healthy samples, aimed at facilitating automated disease diagnosis using computer vision.
Reproduction assets foundThe paper's core asset is its own tomato leaf image dataset (731 raw images plus annotations), publicly deposited on Mendeley Data with an explicit direct URL and DOI.
Dataset · publicr types. Data source location Tomato is one of the most commonly cultivated vegetables in Bangladesh. We collected our data from these three locations: 1. Dinajpur 2. Thakurgaon 3. Kushtia Data accessibility The dataset is published in Mendeley Data. • Data identification number(doi): 10.17632/bpfd9cns5g.2 • Direct URL to data: https://data.mendeley.com/datasets/bpfd9cns5g/2 1 Value of the Data The dataset is highly valuable for agricultural research and machine learning applications, especially in tomato cultivation, providing valuable insights and potential impacts. • The Tomato Leaf Dataset [ 2 ] provides a comprehensive collection of tomato leaf, images, categorized by disease type, aidiOpen asset ↗Mendeley Data · 10.17632/bpfd9cns5g.2lines:1-62
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published22 Mar 2025Data in briefCited by 8 · OpenAlex ↗

OliveTreeCrownsDb: A high-resolution UAV dataset for detection and segmentation in agricultural computer vision.

OliveAerial / UAVWhole plant / canopy / plot / fieldObject detectionSegmentation

This article introduces OliveTreeCrownsDb, a comprehensive dataset of high-resolution images captured by a DJI Phantom 4 RTK drone. The dataset includes 46 images covering an entire olive farm, focusing on the detection and analysis of olive tree crowns and supporting segmentation tasks. Each image is accompanied by detailed metadata, such as focal distance, capture altitude, GPS coordinates, and other essential parameters for accurate tree mapping and localization. OliveTreeCrownsDb is publicly accessible, promoting research in precision agriculture, including tree crown detection, segmentation, geometric shape analysis, automation, yield estimation, and computer vision applications. It facilitates the development of innovative algorithms to optimize resource allocation and improve crop management. By enabling studies on tree crown analysis and farm monitoring, OliveTreeCrownsDb advances agricultural technologies and enhances management practices in olive cultivation.

Why it matches plant phenotyping methodsオリーブ樹冠の画像検出・セグメンテーションと幾何形状解析を可能にする公開データセットが研究の中心であり、植物の樹冠形態を抽出する再利用可能な基盤に該当する。

abstractThis article introduces OliveTreeCrownsDb, a comprehensive dataset of high-resolution images captured by a DJI Phantom 4 RTK drone.
Reproduction assets foundThe paper's own UAV olive tree crown dataset (images, annotations, point cloud, DEM) is publicly deposited on Mendeley Data with explicit direct URL and DOI.
Dataset · public/ Town / Region: Meknas farm site Country: Morocco The GPS coordinates of the olive farm are 33°53′17"N 5°25′22"W, or in decimal format: 33.88802°N, -5.42281°W. Data accessibility Repository name: OliveTreeCrownsDb Data identification number : doi: 10.17632/xym8rd2srf.2 Direct URL to data: Instructions for accessing these data: https://data.mendeley.com/datasets/xym8rd2srf/2 Related research article none 1. Value of the Data The OliveTreeCrownsDb dataset is a valuable resource for research in computer vision and precision agriculture. Here are the key aspects that highlight its importance: • Unique and Specialized Source: OliveTreeCrownsDb offers an exclusive high-resolution dataset specificOpen asset ↗10.17632/xym8rd2srf.2lines:1-54
Code / dataset availability confirmedarXiv · checked 13 Sept 2026
Published17 Mar 2025arXiv

3D Hierarchical Panoptic Segmentation in Real Orchard Environments Across Different Sensors

AppleField / plotLiDAR / point cloudRGB-D / ToFFruitStem / branchWhole plant / canopy / plot / fieldCountingSegmentation

Crop yield estimation is a relevant problem in agriculture, because an accurate yield estimate can support farmers' decisions on harvesting or precision intervention. Robots can help to automate this process. To do so, they need to be able to perceive the surrounding environment to identify target objects such as trees and plants. In this paper, we introduce a novel approach to address the problem of hierarchical panoptic segmentation of apple orchards on 3D data from different sensors. Our approach is able to simultaneously provide semantic segmentation, instance segmentation of trunks and fruits, and instance segmentation of trees (a trunk with its fruits). This allows us to identify relevant information such as individual plants, fruits, and trunks, and capture the relationship among them, such as precisely estimate the number of fruits associated to each tree in an orchard. To efficiently evaluate our approach for hierarchical panoptic segmentation, we provide a dataset designed specifically for this task. Our dataset is recorded in Bonn, Germany, in a real apple orchard with a variety of sensors, spanning from a terrestrial laser scanner to a RGB-D camera mounted on different robots platforms. The experiments show that our approach surpasses state-of-the-art approaches in 3D panoptic segmentation in the agricultural domain, while also providing full hierarchical panoptic segmentation. Our dataset is publicly available at https://www.ipb.uni-bonn.de/data/hops/. The open-source implementation of our approach is available at https://github.com/PRBonn/hapt3D.

Why it matches plant phenotyping methodsリンゴ樹・果実・幹を3Dセグメンテーションし、樹ごとの果実数を推定する手法と専用データセットを中心に開発・評価しており、植物の器官形態・収量関連形質の取得に該当する。

abstractwe introduce a novel approach to address the problem of hierarchical panoptic segmentation of apple orchards on 3D data from different sensors.
Reproduction assets foundThe paper introduces the HOPS dataset of annotated 3D apple orchard point clouds (TLS, UAV, UGV, SfM) for hierarchical panoptic segmentation, publicly available at the authors' IPB Bonn page, and releases the open-source implementation (hapt3D) on GitHub. Both are paper-specific, public, and actionable.
Code · publicThe open-source implementation of our approach is available at https://github.com/PRBonn/hapt3D .Open asset ↗PRBonn/hapt3Dlines:1-59
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published13 Mar 2025iMetaCited by 10 · OpenAlex ↗

Phenotyping, genome-wide dissection, and prediction of maize root architecture for temperate adaptability.

MaizeMorphology / geometry measurementRoot system architecture

Abstract Root System Architecture (RSA) plays an essential role in influencing maize yield by enhancing anchorage and nutrient uptake. Analyzing maize RSA dynamics holds potential for ideotype‐based breeding and prediction, given the limited understanding of the genetic basis of RSA in maize. Here, we obtained 16 root morphology‐related traits (R‐traits), 7 weight‐related traits (W‐traits), and 108 slice‐related microphenotypic traits (S‐traits) from the meristem, elongation, and mature zones by cross‐sectioning primary, crown, and lateral roots from 316 maize lines. Significant differences were observed in some root traits between tropical/subtropical and temperate lines, such as primary and total root diameters, root lengths, and root area. Additionally, root anatomy data were integrated with genome‐wide association study (GWAS) to elucidate the genetic architecture of complex root traits. GWAS identified 809 genes associated with R‐traits, 261 genes linked to W‐traits, and 2577 key genes related to 108 slice‐related traits. We confirm the function of a candidate gene, fucosyltransferase5 ( FUT5 ), in regulating root development and heat tolerance in maize. The different FUT5 haplotypes found in tropical/subtropical and temperate lines are associated with primary root features and hold promising applications in molecular breeding. Furthermore, we performed machine learning prediction models of RSA using root slice traits, achieving high prediction accuracy. Collectively, our study offers a valuable tool for dissecting the genetic architecture of RSA, along with resources and predictive models beneficial for molecular design breeding and genetic enhancement.

Why it matches plant phenotyping methodsトウモロコシ根系形態を大規模に取得し、根スライス形質に基づく機械学習予測モデルと再利用可能な資源を構築しており、表現型取得・推定が研究の主要部分です。

abstractwe obtained 16 root morphology‐related traits (R‐traits), 7 weight‐related traits (W‐traits), and 108 slice‐related microphenotypic traits (S‐traits)
Reproduction assets foundThe paper's root phenotyping images, phenotypic data, and RNA-seq data are deposited on figshare, and the authors' GWAS analysis pipeline code is publicly available on GitHub, both explicitly stated in the Data Availability Statement.
Dataset · publicAll the images, phenotypic data, and RNAs‐seq data are available at https://doi.org/10.6084/m9.figshare.27605208.v1 .Open asset ↗figshare · 10.6084/m9.figshare.27605208.v1lines:197-303
Code · publicThe original data and code for GWAS analysis pipelines can be downloaded at https://github.com/GUOWEIJUN/maizerootphenomics .Open asset ↗GitHub · GUOWEIJUN/maizerootphenomicslines:197-303
Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Published10 Mar 2025AgricultureCited by 24 · OpenAlex ↗

Plant Disease Segmentation Networks for Fast Automatic Severity Estimation Under Natural Field Scenarios

AppleSoybeanWheatField / plotLaboratory / benchtopLeafWhole plant / canopy / plot / fieldSegmentationStress / disease detectionDisease symptoms / severity

The segmentation of plant disease images enables researchers to quantify the proportion of disease spots on leaves, known as disease severity. Current deep learning methods predominantly focus on single diseases, simple lesions, or laboratory-controlled environments. In this study, we established and publicly released image datasets of field scenarios for three diseases: soybean bacterial blight (SBB), wheat stripe rust (WSR), and cedar apple rust (CAR). We developed Plant Disease Segmentation Networks (PDSNets) based on LinkNet with ResNet-18 as the encoder, including three versions: ×1.0, ×0.75, and ×0.5. The ×1.0 version incorporates a 4 × 4 embedding layer to enhance prediction speed, while versions ×0.75 and ×0.5 are lightweight variants with reduced channel numbers within the same architecture. Their parameter counts are 11.53 M, 6.50 M, and 2.90 M, respectively. PDSNetx0.5 achieved an overall F1 score of 91.96%, an Intersection over Union (IoU) of 85.85% for segmentation, and a coefficient of determination (R2) of 0.908 for severity estimation. On a local central processing unit (CPU), PDSNetx0.5 demonstrated a prediction speed of 34.18 images (640 × 640 pixels) per second, which is 2.66 times faster than LinkNet. Our work provides an efficient and automated approach for assessing plant disease severity in field scenarios.

Why it matches plant phenotyping methods植物病害画像から病斑割合と病害重症度を推定する画像セグメンテーション手法を開発し、野外データセット、精度、速度を評価しており、植物表現型取得法が中心である。

abstractThe segmentation of plant disease images enables researchers to quantify the proportion of disease spots on leaves, known as disease severity.
Reproduction assets foundThe paper's field-scenario plant disease image dataset (SBB, WSR, CAR with three-color pixel labels) is publicly released on Kaggle via DOI, as stated in the Data Availability Statement. No author analysis code or trained model checkpoints are explicitly deposited.
Dataset · publicData Availability Statement: The original data presented in this study are openly available in Kaggle at https://doi.org/10.34740/kaggle/ds/6620728, accessed on 9 March 2025.Open asset ↗Kaggle · 10.34740/kaggle/ds/6620728pdf-page:15 lines:1-58
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published7 Mar 2025Scientific dataCited by 2 · OpenAlex ↗

Fire ecology database for documenting plant responses to fire events in Australia.

Field / plotWhole plant / canopy / plot / fieldVisualization / data managementStress response / tolerance

An understanding of fire-response traits is essential for predicting how fire regimes structure plant communities and for informing fire management strategies for biodiversity conservation. Quantification of these traits is complex, encompassing several levels of data abstraction scaling up from field observations of individuals, to general categories of species responses. We developed the Fire Ecology Database to accommodate this complexity. Its conceptual framework is underpinned by a flexible data pipeline enabling links between fire-related trait data and event information at individual, population, and community levels. Key features include: (a) concise and documented trait and method vocabularies; (b) documented uncertainty in observations and aggregation; and (c) documented origin of data including field observations, laboratory experiments, and expert elicitation. We demonstrated application of our framework using data from new field surveys and existing data sets in New South Wales, Australia. The database includes 14 traits for 6,287 plant species derived from 8,936 field work records from 2007 to 2018, 7,054 field records from surveys after 2019, and 48,306 records from 301 existing sources.

Why it matches plant phenotyping methods火災応答形質を体系的に収集・標準化するデータベースとデータパイプライン自体が中心的な方法論的貢献であり、植物形質データの不確実性・測定法・由来も記録しているため、フェノタイピング用データ基盤として含める。

abstractWe developed the Fire Ecology Database to accommodate this complexity. Its conceptual framework is underpinned by a flexible data pipeline enabling links between fire-related trait data and event information at individual, population, and community levels.
Reproduction assets foundThe paper's core outputs (Fire Ecology Database v1.1 SQL dump, R data frames, CSV/XLSX exports on FigShare/OSF, and the Python import scripts/Jupyter notebooks) are stated to be publicly available, but no concrete repository URL or identifier for them appears in the supplied blocks, and none matches an allowed URL, so
Code · publicCustomised scripts were written in Python to automate the importation of field data from the spreadsheets into the database. These scripts are available for download (see Code availability section)Open asset ↗pdf-page:6 lines:1-78
Dataset · publicStatic versions of the Fire Ecology Database, including version 1.1 used in this descriptor, are available via FigShare or OSF in three different formatsOpen asset ↗FigSharepdf-page:9 lines:1-78
Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Published4 Mar 2025Remote SensingCited by 7 · OpenAlex ↗

Improved Detection and Location of Small Crop Organs by Fusing UAV Orthophoto Maps and Raw Images

Aerial / UAVField / plotWhole plant / canopy / plot / fieldAnnotation / quality controlObject detection

Extracting the quantity and geolocation data of small objects at the organ level via large-scale aerial drone monitoring is both essential and challenging for precision agriculture. The quality of reconstructed digital orthophoto maps (DOMs) often suffers from seamline distortion and ghost effects, making it difficult to meet the requirements for organ-level detection. While raw images do not exhibit these issues, they pose challenges in accurately obtaining the geolocation data of detected small objects. The detection of small objects was improved in this study through the fusion of orthophoto maps with raw images using the EasyIDP tool, thereby establishing a mapping relationship from the raw images to geolocation data. Small object detection was conducted by using the Slicing-Aided Hyper Inference (SAHI) framework and YOLOv10n on raw images to accelerate the inferencing speed for large-scale farmland. As a result, comparing detection directly using a DOM, the speed of detection was accelerated and the accuracy was improved. The proposed SAHI-YOLOv10n achieved precision and mean average precision (mAP) scores of 0.825 and 0.864, respectively. It also achieved a processing latency of 1.84 milliseconds on 640×640 resolution frames for large-scale application. Subsequently, a novel crop canopy organ-level object detection dataset (CCOD-Dataset) was created via interactive annotation with SAHI-YOLOv10n, featuring 3986 images and 410,910 annotated boxes. The proposed fusion method demonstrated feasibility for detecting small objects at the organ level in three large-scale in-field farmlands, potentially benefiting future wide-range applications.

Why it matches plant phenotyping methodsUAV画像と生画像の融合、SAHI-YOLOv10nによる作物器官の検出・位置推定を中心に開発・評価し、器官レベルの大規模データセットも構築しているため、植物表現型取得法が研究の中心である。

abstractThe detection of small objects was improved in this study through the fusion of orthophoto maps with raw images using the EasyIDP tool, thereby establishing a mapping relationship from the raw images to geolocation data.
Reproduction assets foundThe paper's CCOD-Dataset (3986 UAV images, 410,910 annotated bounding boxes of crop canopy organs) is publicly released on Hugging Face, and the authors' SAHI-YOLOv10 detection framework code is hosted in a public GitHub repository. Other code (fusion/EasyIDP pipeline) is only available upon request.
Dataset · publicThe CCOD-Dataset link is publicly available at Hugging Face at https://huggingface.co/datasets/Nir-Open asset ↗CCOD-Datasetpdf-page:8 lines:1-51
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 6 Sept 2026
Published1 Mar 2025Plant PhenomicsCited by 11 · OpenAlex ↗

CVRP: A rice image dataset with high-quality annotations for image segmentation and plant phenomics research.

RiceField / plotLaboratory / benchtopPanicle / ear / spikeSeed / grainWhole plant / canopy / plot / fieldCounting2D/3D reconstructionSegmentation

Machine learning models for crop image analysis and phenomics are highly important for precision agriculture and breeding and have been the subject of intensive research. However, the lack of publicly available high-quality image datasets with detailed annotations has severely hindered the development of these models. In this work, we present a comprehensive multicultivar and multiview rice plant image dataset (CVRP) created from 231 landraces and 50 modern cultivars grown under dense planting in paddy fields. The dataset includes images capturing rice plants in their natural environment, as well as indoor images focusing specifically on panicles, allowing for a detailed investigation of cultivar-specific differences. A semiautomatic annotation process using deep learning models was designed for annotations, followed by rigorous manual curation. We demonstrated the utility of the CVRP by evaluating the performance of four state-of-the-art (SOTA) semantic segmentation models. We also conducted 3D plant reconstruction with organ segmentation via images and annotations. The database not only facilitates general-purpose image-based panicle identification and segmentation but also provides valuable resources for challenging tasks such as automatic rice cultivar identification, panicle and grain counting, and 3D plant reconstruction. The database and the model for image annotation are available at https://bic.njau.edu.cn/CVRP.html.

Why it matches plant phenotyping methodsイネ画像データセットとアノテーションモデルを開発・評価し、セグメンテーション、器官再構成、穂・粒数計測などの再利用可能な表現型解析を中心に扱っているため。

abstractwe present a comprehensive multicultivar and multiview rice plant image dataset (CVRP)
Reproduction assets foundThe paper's own CVRP rice image dataset (images + annotations), accompanying code, and trained Mask2Former annotation model are explicitly stated as publicly available on Hugging Face and the authors' NJAU site.
Dataset · publicThe CVRP dataset is publicly available on Hugging Face at https://huggingface.co/datasets/CVRPDataset/CVRP for academic use under the specified license.Open asset ↗CVRPDataset/CVRPhtml-lines:236-252
Code · publicThe accompanying code and trained models are available at https://huggingface.co/CVRPDataset/Model.Open asset ↗CVRPDataset/Modelhtml-lines:236-252
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 Feb 2025Data in briefCited by 5 · OpenAlex ↗

A data-driven approach to turmeric disease detection: Dataset for plant condition classification.

LeafRootClassificationDisease symptoms / severity

Turmeric, Curcuma longa, is an economically and medicinally important crop. However, the crop has often suffered from diseases such as rhizome disease roots, leaf blotch, and dry conditions of leaves. The control of these diseases essentially requires early and accurate diagnosis to reduce losses and help farmers adopt sustainable farming methods. The conventional methods of diagnosis involve a visual examination of symptoms, which is laborious, subjective, and rather impossible in large areas. This paper proposes a new dataset consisting of 1037 originals and 4628 augmented images of turmeric plants representing five classes: healthy leaf, dry leaf, leaf blotch, rhizome disease roots, and rhizome healthy roots. The dataset was pre-processed to enhance its applicability to deep learning applications by resizing, cleaning, and augmenting the data through flipping, rotation, and brightness adjustment. The turmeric plant disease classification was conducted using the Inception-v3 model, attaining an accuracy of 97.36% with data augmentation, compared to 95.71% without augmentation. Some of the major key performance metrics are precision, recall, and F1-score, which establish the efficacy and robustness of the model. This work attempts to show the potential of AI-aided solutions towards precision farming and sustainable crop production in developing agriculture disease management. The publicly available dataset and the results obtained are expected to attract more research interest for innovations in AI-driven agriculture .

Why it matches plant phenotyping methodsウコン植物の葉・根の病徴を画像から分類する公開データセットと解析手法を構築・評価しており、植物の病害状態を推定するフェノタイピング手法が中心である。

abstractThis paper proposes a new dataset consisting of 1037 originals and 4628 augmented images of turmeric plants representing five classes: healthy leaf, dry leaf, leaf blotch, rhizome disease roots, and rhizome healthy roots.
Reproduction assets foundThe paper's turmeric plant disease image dataset (1073 original + 4628 augmented images, five classes) is publicly deposited on Mendeley Data with an explicit DOI and direct URL. No separate analysis code repository is stated.
Dataset · publicels in the early detection and effective management of diseases affecting turmeric plants to support sustainable agriculture. Data source location Town/City/Region: Charpolisha, Jamalpur Country: Bangladesh . Data accessibility Repository name: Mendeley Data. Data identification number: 10.17632/g46dvrcvwn.2 Direct URL to data: https://data.mendeley.com/datasets/g46dvrcvwn/2 Related research article None . 1. Value of the Data • This dataset consists of several images regarding turmeric plant diseases, starting from the most prevalent to the rarest. Hence, it is quite valuable in terms of scientific research and agriculture. This will act as a stepping stone to further improve the plant pathOpen asset ↗Mendeley Data · 10.17632/g46dvrcvwn.2lines:1-46
Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Published24 Feb 2025Plant MethodsCited by 18 · OpenAlex ↗

A method for phenotyping lettuce volume and structure from 3D images

LettuceLiDAR / point cloudRGB-D / ToFWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionYield / biomass estimationArchitecture / morphology / geometryBiomass / plant weightLeaf traits

Abstract Monitoring plant growth is crucial for effective crop management, and using color and depth (RGBD) cameras to model lettuce has emerged as one of the most convenient and non-invasive methods. In recent years, deep learning techniques, particularly neural networks, have become popular for estimating lettuce fresh weight. However, these models are typically specific to particular datasets, lack domain adaptation, and are often limited by the availability of open-access datasets. In this study, we propose a method based on plant geometric features for estimating the rosette structure and volume of lettuce. This new approach was compared to existing methods that reconstruct surfaces from point clouds, such as Ball Pivoting and Alpha Shapes. The proposed method creates a tight hull around the plant's point cloud, preserving high detail of the rosette structure while filling in surface holes in areas not visible to 3D cameras. Using a linear regression model, we estimated fresh weight for this dataset, achieving a root mean square error (RMSE) of 18.2 g when using only the estimated plant volume, and 17.3 g when both volume and geometric features were included. Additionally, we introduced new geometric features that characterize leaf density, which could be useful for breeding applications. A dataset of 402 point clouds of lettuce plants, captured before harvest, was compiled using one top-down and three side-view 3D cameras.

Why it matches plant phenotyping methodsRGB-D画像からレタスの構造・体積・葉密度を抽出し、生体重推定を検証する手法開発が研究の中心であり、データセットも構築している。

abstractIn this study, we propose a method based on plant geometric features for estimating the rosette structure and volume of lettuce.
Reproduction assets foundThe paper's own lettuce 3D point cloud dataset (Pii, 402 point clouds with fresh weight references) is deposited on Zenodo, and the vacuum-package surface reconstruction code plus data processing scripts are publicly available on the authors' GitHub repository. Both are paper-specific, public, and actionable.
Dataset · publicData used in this study and developed models are available on Zenodo storage service https://zenodo.org/records/8410252 .Open asset ↗Zenodo · 8410252lines:158-220
Code · publicThe code used at this study is available at https://github.com/VicB18/LettuceFW (accessed on 1 November 2024).Open asset ↗GitHub · VicB18/LettuceFWlines:158-220
Code · publicThe code for the vacuum package method, along with the data processing scripts used in this study, are available in the Supplementary Information.Open asset ↗lines:98-114
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published21 Feb 2025Data in briefCited by 3 · OpenAlex ↗

Colombian coffee tree leaves multispectral images dataset.

CoffeeField / plotRGB / grayscaleMultispectral / hyperspectralLeafDisease symptoms / severity

In this work, a unique database of 6726 multispectral images of coffee leaves is presented. These images were captured in JPG format for the RGB photos and in TIF format for the five multispectral bands: blue, green, red, NIR and red edge, providing a detailed view of different wavelengths of the electromagnetic spectrum. Images in TIF format have a color depth of 16 bits per pixel, ensuring good quality. The blue band (Band 1) captures light in the blue region of the spectrum, approximately 450 to 500 nm. The green band (Band 2) records light in the green region, approximately between 500 and 620 nm. The red band (Band 3) captures light in the red region, between 620 and 750 nm. The red-edge band (Band 4) lies between the red band and the NIR, and is sensitive to the transition between green vegetation and non-vegetation, around 840 nm. Finally, the near infrared band (Band 5) captures light in the near infrared region, between 750 and 900 nm. For ease of identification, images are labeled as follows: if the image name ends in 0, it is an RGB image; if it ends in 1, it corresponds to the blue band; if it ends in 2, to the green band; if it ends in 3, to the red band; if it ends in 4, to the red-edge band; and if it ends in 5, to the near-infrared band. The images show coffee leaves with and without lesions caused by the Hemileia vastatrix fungus, known as coffee rust. These samples were collected from Colombian coffee farms and the images were captured under controlled lighting conditions to ensure quality and consistency. This database is an invaluable resource for precision agriculture research and early detection of crop diseases. With these 6726 images, researchers can use advanced image processing and machine learning techniques to identify differences between healthy leaves and those affected by rust. This can lead to the development of effective predictive models, enabling early detection and more efficient management of diseases in coffee plantations, optimizing production and reducing economic losses for farmers.

Why it matches plant phenotyping methodsコーヒー葉の病斑という植物の病害状態を対象としたマルチスペクトル画像データセットであり、再利用可能なフェノタイピング用データセットの提供が中心です。

abstractIn this work, a unique database of 6726 multispectral images of coffee leaves is presented.
Reproduction assets foundThe paper is a data descriptor whose own multispectral coffee leaf image dataset is publicly deposited on Kaggle with an explicit direct URL and DOI, matching an allowed URL.
Dataset · publicth of 16 bits per pixel . Data source location Institution: Escuela Colombiana de Ingeniería Julio Garavito University City/Town/Region: Bogotá D.C. Country: Colombia Latitude: 4.5983° * Longitude: 74.0051°. Data accessibility Repository name: Coffe Rust Data identification number: 10.34740/kaggle/ds/5644659 Direct URL to data: https://www.kaggle.com/ds/5644659 Instructions for accessing these data: Data available free of charge to anyone with access to the Internet and the web server address provided. Related research article [ 1 ] Jorge Luis Aroca Trujillo, Alexander Pérez-Ruiz. “Technologies Applied in the Field of Early Detection of Coffee Rust Fungus Diseases: A Review.” Nongye JOpen asset ↗Kaggle · 10.34740/kaggle/ds/5644659lines:1-53
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published21 Feb 2025Data in briefCited by 2 · OpenAlex ↗

Drone-based dataset of annotated sunflower images from Bangladesh.

SunflowerAerial / UAVField / plotWhole plant / canopy / plot / fieldClassificationObject detectionGrowth / development / phenologyStress response / tolerance

Accurate and automated detection of sunflower plants, along with assessments of their growth stages and health conditions, is crucial for enabling precision agriculture and improving crop management. In this work, we present a drone-based dataset of annotated sunflower images, derived from high-resolution videos captured at two distinct locations in Bangladesh. The original dataset comprises 1649 images extracted from drone footage of the BARI Surjomukhi-3 variety under various orientations, health conditions, and weather scenarios. After meticulous annotation using the Roboflow platform and augmentation with seven distinct techniques, the dataset expanded to 4286 images in Pascal VOC format. Detailed metadata-including geospatial coordinates, timestamped acquisition conditions, and camera settings-accompanies the dataset to support reproducibility and model generalization. By offering a comprehensive suite of annotated and augmented images, this dataset provides a valuable resource for developing and refining computer vision models geared toward sunflower detection, maturity evaluation, and yield prediction, ultimately advancing sustainable farming practices and decision-making tools in agricultural research.

Why it matches plant phenotyping methodsヒマワリ画像を注釈付きデータセットとして構築し、成長段階・健康状態・成熟度などの植物状態推定を支援することが中心であり、再利用可能な画像ベース表現型データセットに該当する。

abstractwe present a drone-based dataset of annotated sunflower images
Reproduction assets foundThis Data in Brief article describes a drone-based annotated sunflower image dataset from Bangladesh, publicly deposited on Mendeley Data (DOI 10.17632/txct4k36ct.1) with a companion Roboflow Universe project for annotation conversion. Both are paper-specific, public, and directly actionable.
Dataset · publicrsingdi, and Amjhupi, Meherpur Country: Bangladesh Latitude and longitude: Nagoriakandi, Narsingdi: Latitude 23.906801° N, Longitude 90.710563° E Amjhupi, Meherpur: Latitude 23.744897° N, Longitude 88.69174° E Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/txct4k36ct.1 Direct URL to data: https://data.mendeley.com/datasets/txct4k36ct/1 1. Value of the Data • Drone-captured, high-resolution images meticulously annotated for sunflower detection, growth stage, and health conditions enable the development of precise computer vision models [ 1 ]. • Unlike conventional drone images taken from overhead perspectives, the dataset includes images captured at lowOpen asset ↗Mendeley Data · 10.17632/txct4k36ct.1lines:1-51
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published18 Feb 2025Data in briefCited by 4 · OpenAlex ↗

Early detection of Zymoseptoria tritici infection on wheat leaves using hyperspectral imaging data.

WheatMultispectral / hyperspectralLeafStress / disease detectionDisease symptoms / severity

This article presents a hyperspectral imaging (HSI) database of healthy leaves and leaves infected with Zymoseptoria tritici fungal pathogen responsible for leaf blotch (Lb) disease. Leaves of two durum wheat genotypes were studied under controlled conditions to track the evolution of Lb disease and capture significant spectral and spatial differences until the onset of symptoms. Hyperspectral image acquisitions were purchased with two cameras in visible-near infrared (VNIR) and short-wave infrared (SWIR) spectral ranges on eighteen dates between one day before inoculation and twenty days after inoculation. For each wavelength range studied, a total of 1175 images provided information on 3326 leaves measured throughout the experiment. These data are valuable since they can be used as a basis to monitor disease's development over time, to build leaf classification models according to their infection status per genotype per day, to develop prediction models related to symptoms' appearance, or to test imaging and spectral analysis methods.

Why it matches plant phenotyping methodsコムギ葉の病害状態をハイパースペクトル画像で取得したデータベースを構築し、感染状態分類・症状出現予測や画像解析手法の評価基盤として提供しており、表現型取得法が中心である。

abstractThis article presents a hyperspectral imaging (HSI) database of healthy leaves and leaves infected with Zymoseptoria tritici fungal pathogen responsible for leaf blotch (Lb) disease.
Reproduction assets foundThe paper is a Data in Brief article describing a public hyperspectral imaging dataset of healthy and Zymoseptoria tritici-infected durum wheat leaves, deposited on Data INRAE with DOI 10.57745/WVP0FJ. This is the paper's own plant-phenotyping measurement data (VNIR/SWIR hyperspectral images, pixel coordinates, and CSV
Dataset · publicand HySpex SWIR-384 (Norsk Elektro Optikk, Norway). Data source location Institution: Institut National de Recherche pour l'Agriculture, l'Alimentation et l'Environnement (INRAE) City: Montpellier Country: France Data accessibility Repository name: Data INRAE Data identification number: doi: 10.57745/WVP0FJ Direct URL to data: https://doi.org/10.57745/WVP0FJ 1 Value of the Data • This dataset depicts the visual appearance and spectral information related to the onset kinetics of Lb disease symptoms on wheat leaves using hyperspectral images acquired post-inoculation. • The images captured are valuable to monitor the evolution of the Lb disease on wheat leaves through the developmenOpen asset ↗Data INRAE · 10.57745/WVP0FJlines:1-60
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published14 Feb 2025Data in briefCited by 3 · OpenAlex ↗

Towards precision agriculture: A dataset for early detection of corn leaf pests.

MaizeField / plotLeafClassificationSegmentationDisease symptoms / severity

Corn ( Zea mays ), commonly referred to as Indian wheat, is a widely cultivated tropical annual herbaceous plant of the Poaceae family. It is primarily grown for its starch-rich grains and as a forage crop. In Cameroon, corn is the most consumed cereal, surpassing rice and sorghum, with an estimated production of 2.2 million tons annually. However, corn production is frequently threatened by insect infestations, which hinder crop development, reduce yields, and degrade its quality. Early detection of insect attacks is essential for farmers, as timely intervention can prevent widespread damage, reduce pesticide usage, and improve production yields. Insect infestations on corn manifest through various symptoms on leaves, stems, and seeds. Among these, foliar attacks are particularly detrimental, disrupting plant growth and significantly reducing yields. Symptoms of these attacks include leaf perforations, yellowing, and white spot deposits, ultimately altering the leaf texture. To address these challenges, machine learning models offer a promising solution for early detection of foliar attacks, enabling farmers to take timely and effective action. This paper introduces a dataset focused on three major pests: Spodoptera frugiperda (Fall Armyworm), Helminthosporium leaf blight, and Zonocerus variegatus (Variegated Grasshopper), which are among the most frequent and destructive agents affecting corn crops. The dataset comprises images of corn leaves captured in natural environments at various growth stages and field locations. Images were taken using smartphone cameras at different times of the day, providing diverse lighting conditions, and in various fields, which introduced several background contaminations, ensuring a realistic representation of field conditions. The dataset comprises eight directories: two containing healthy leaf images (1308 without augmentation and 11,772 with augmentation), two containing manually segmented backgrounds of healthy leaves (1308 without augmentation and 11,772 with augmentation), two containing healthy leaves with CNDVI algorithm-segmented backgrounds (1308 without augmentation and 11,772 with augmentation), one containing 848 infected images with manually segmented backgrounds and highlighted infected areas, and one containing 7632 augmented versions of the infected images. This dataset serves as a valuable resource for researchers and students, providing opportunities to develop machine learning and deep learning models for corn disease detection, classification, natural image segmentation, and model interpretability and explainability. By facilitating advancements in precision agriculture and automated pest detection, the dataset contributes to sustainable agricultural practices and the broader field of agroinformatics.

Why it matches plant phenotyping methodsトウモロコシ葉の病害・害虫症状を画像化し、セグメンテーション済みデータセットとして提供することが中心で、植物の病害状態を直接推定する画像ベースの表現型手法に該当する。

abstractThis paper introduces a dataset focused on three major pests: Spodoptera frugiperda (Fall Armyworm), Helminthosporium leaf blight, and Zonocerus variegatus (Variegated Grasshopper)
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository of corn leaf pest images (healthy and infected, with manual and CNDVI-based segmentation and annotations), directly reproducing this paper's phenotyping measurements. The dataset URL is explicitly given and matches an allowed URL.
Dataset · publicogbessou PK17 (Latitude: 4.100316; Longitude: 9.802564) • Papas (Latitude: 4.057849; Longitude: 9.819752) 3. On village of Moungo division in littoral region: Edjocmoa (Latitude: 4.980205; Longitude: 9.946348) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/ymvghfcww7.1 Direct URL to data: https://data.mendeley.com/datasets/ymvghfcww7/1 Related research article A robust segmentation method combined with classification algorithms for field-based diagnosis of maize plant phytosanitary state [ 1 ] . 1. Value of the Data • The data facilitate early detection and monitoring of major corn leaf diseases. The dataset enables the early identification and continuOpen asset ↗Mendeley Data · 10.17632/ymvghfcww7.1lines:38-76
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published12 Feb 2025Data in briefCited by 4 · OpenAlex ↗

High-resolution dataset for tea garden disease management: Precision agriculture insights.

TeaField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

The economic development of many countries largely depends on tea plantations that suffer from diseases adversely affecting their productivity and quality. This study presents a high-resolution dataset aimed at advancing precision agriculture for managing tea garden diseases. The size of the dataset is 3960 images and pixel dimension is (1024 × 1024) of the images were collected by using smartphones. This dataset contains detailed images of Tea Leaf Blight, Tea Red Leaf Spot and Tea Red Scab maladies inflicted on tea leaves as well as environmental statistics and plant health. The images were captured and stored in JPG format. The main aim of this dataset is to provide tool for detection and classification of different types of tea garden disease. Applying this dataset will enable the development of early detection systems, best-practice care regimens, and enhanced general garden upkeep. A range of images presenting the most prevalent diseases afflicting tea plants are paired with images of healthy leaves to provide a comprehensive overview of all the circumstances that can arise in a tea plantation. Therefore, it can be used to automate diseases tracking, targeted pesticide spraying, and even the making of smart farm tools with development of smart agricultural tools hence enhancing sustainability and efficiency in tea production. This dataset not only provides a strong foundation for applying precision techniques in tea cultivation in agriculture, but also can become an invaluable asset to scientists studying the issues of tea production.

Why it matches plant phenotyping methods茶葉の病害状態を画像で記録したデータセットであり、植物病害の画像ベース表現型評価を支えるデータ資源が中心です。

abstractThis study presents a high-resolution dataset aimed at advancing precision agriculture for managing tea garden diseases.
Reproduction assets foundThe paper is a Data in Brief article describing a tea leaf disease image dataset (3960 original images, 4000 augmented) publicly deposited on Mendeley Data with an explicit DOI and direct URL. This is a paper-specific, public, actionable plant-phenotyping asset (plant images used for disease classification phenotyping)
Dataset · publicar, Sylhet, Bangladesh. The project was conducted under the supervision of an expert from Bangladesh's Ministry of Agriculture. Data source location Location: Moulvi Bazar tea garden,Sylhet Country: Bangladesh Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/tt2smzrzrs.4 Direct URL to data: https://data.mendeley.com/datasets/tt2smzrzrs/4 Related research article None 1 Value of the Data • Tea is a major global agricultural crop with economic implications as well as cultural significance. This drink is famous for diverse tastes and health benefits. In many civilizations, tea remains their main beverage [ 1 ]. There are several countries that supply most oOpen asset ↗Mendeley Data · 10.17632/tt2smzrzrs.4lines:1-51
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 14 Sept 2026
Published11 Feb 2025Plant MethodsCited by 3 · OpenAlex ↗

Deep-learning-ready RGB-depth images of seedling development.

RGB-D / ToFWhole plant / canopy / plot / fieldAnnotation / quality controlGrowth / time-series analysisGrowth / development / phenology

In the era of machine learning-driven plant imaging, the production of annotated datasets is a very important contribution. In this data paper, a unique annotated dataset of seedling emergence kinetics is proposed. It is composed of almost 70,000 RGB-depth frames and more than 700,000 plant annotations. The dataset is shown valuable for training deep learning models and performing high-throughput phenotyping by imaging. The ability of such models to generalize to several species and outperform the state-of-the-art owing to the delivered dataset is demonstrated. We also discuss how this dataset raises new questions in plant phenotyping.

Why it matches plant phenotyping methods植物の出芽速度を対象とする大規模RGB深度画像・アノテーションデータセットを提供し、深層学習および高スループット表現型解析への利用性を実証しており、表現型取得基盤が中心である。

abstracta unique annotated dataset of seedling emergence kinetics is proposed
Reproduction assets foundThis is a data paper whose core contribution is a public annotated RGB-depth seedling dataset (~70,000 frames, >700,000 annotations) deposited in DATA INRAE with DOI 10.57745/AMFJTK, explicitly stated as publicly accessible. Other allowed URLs (license, Intel datasheet, Jülich record) are not paper-specific assets.
Dataset · publicSynthesis of the full time-lapse and RGB-Depth full frame quantity per species Species Pots time-lapse Labelled pots time-lapse RGB-depth full frame Rapeseed 1 760 336 15 218 Tomatoes 1 960 480 33 283 Beans 2 320 400 21 445 Total 6 040 1 216 69 946 The dataset is publicly accessible in the DATA INRAE repository, DOI: https://doi.org/10.57745/AMFJTK . The file tree structure is illustrated in Fig. 4 . The dataset is organized into 11 compressed .zip files, each corresponding to a distinct trial. Within these files, images are sorted chronologically by acquisition start date, then by camera, and stored in .png format within dedicated color and depth folders. Labels are alsoOpen asset ↗DATA INRAE · 10.57745/AMFJTKlines:105-195
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published7 Feb 2025Data in briefCited by 11 · OpenAlex ↗

Image dataset for classification of diseases in guava fruits and leaves.

FruitLeafClassificationDisease symptoms / severity

Guava (Psidium guajava) this is a tropical fruit and one of the common tropical fruits in Bangladesh. The economic and health value of this important crop is unmeasurable, but it quickly becomes infected with many diseases that can greatly reduce its yield and quality. Thus, the use of technology for automatic fruit and leaf disease detection is necessary in agriculture. The dataset provides us overall guava fruits & leaves image samples for detection of diseases in guava fruits and leaves. This dataset consists of images of healthy and diseased samples infected by fruit disease such as anthracnose, scab, styler end root and leaf disease such as canker, rust, anthracnose and dot. It consists of 3,432 real images obtained from different places in Bangladesh. It is also extended 20,344 augmented images ready to be used for machine learning purposes. This dataset serves as a fundamental building block for utilizing machine learning and computer vision techniques to develop automated detection systems of various diseases. It assists in the early detection of diseases affecting guava, and provides them with solutions to intervene there itself saving agricultural yield and nutritional losses while also promoting sustainable farming practices. This dataset will assist researchers for progressing guava detecting disease through the execution of computational models and application of better machine learning techniques.

Why it matches plant phenotyping methodsグアバ果実・葉の健全/罹病状態を画像で分類するデータセットであり、植物病害状態の画像ベース表現型取得と機械学習利用を中心とするため含める。

titleImage dataset for classification of diseases in guava fruits and leaves.
Reproduction assets foundThe paper is a Data in Brief article describing a public guava fruit/leaf disease image dataset (3,432 original + ~20,344 augmented images) deposited on Mendeley Data with an explicit DOI and direct URL, matching an allowed URL.
Dataset · publicatitude: 23°31′50.2″N , Longitude: 91°05′09.1″E) 2. Asulia Bazar, Dhaka, Bangladesh (Latitude: 23.8971° N , Longitude: 90.3309° E) 3. Baipayl, Dhaka, Bangladesh (Latitude: 23°56′44.0″N Longitude: 90°16′27.3″E) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/fspx44mwfp.1 Direct URL to data: https://data.mendeley.com/datasets/fspx44mwfp/1 1. Value of the Data • The Guava Leaf and Fruit Disease Dataset contains images of healthy guava leaves and fruits and also affected with various diseases such as fruit anthracnose, scab, styler root end, and leaf canker, dot, rust, and anthracnose. This comprehensive dataset allows to develop and train models to detect Open asset ↗Mendeley Data · 10.17632/fspx44mwfp.1lines:1-56
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Feb 2025Data in briefCited by 13 · OpenAlex ↗

TOM2024: Datasets of tomato, onion, and maize images for developing pests and diseases AI-based classification models.

MaizeOnionTomatoField / plotWhole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

The advancement of digital technologies has significantly impacted plant pest and disease management, yet gaps remain, especially in developing regions. This paper introduces the TOM2024 dataset, a comprehensive collection of high-resolution images designed to enhance pest and disease identification of maize, tomato, and onion crops. The dataset encompasses 25,844 raw images and over 12,000 labeled images, categorized into 30 classes (healthy crop, infested crop, and pest) across the three cropping systems. Acquired through meticulous fieldwork in Burkina Faso using high-resolution cameras, the dataset includes diverse environmental conditions and crop stages, ensuring a robust resource for AI model training and validation. The dataset is segmented into three categories: processed images (Category A), selected images with augmentation (Category B), and an online repository with over 25,000 raw images (Category C). Category A and B features images of crops affected by 21 distinct pests and diseases. This dataset addresses critical gaps in existing collections by offering extensive coverage and high-resolution imagery that can be used to developed AI models for automatic identification and classification of pests and diseases that affects crops. TOM2024's versatility extends to research, educational purposes, and the practical application of digital tools in agriculture thereby contributes to the advancement of precision agriculture, sustainable agricultural practices, and food security globally.

Why it matches plant phenotyping methods植物の健全・感染状態を含む画像データセットを構築し、病害・害虫状態の自動分類モデル開発用リソースとして提供することが中心であり、再利用可能な画像ベース表現型データセットに該当する。

abstractThis paper introduces the TOM2024 dataset, a comprehensive collection of high-resolution images designed to enhance pest and disease identification of maize, tomato, and onion crops.
Reproduction assets foundThe paper is a Data in Brief article describing the TOM2024 dataset of tomato, onion, and maize pest/disease images, publicly deposited on Mendeley Data with an explicit direct URL and DOI. This is a paper-specific public image dataset (phenotyping-style plant image asset) directly produced by this paper.
Dataset · publicrce location West African Science Service Centre on Climate Change and Adapted Land Use (WASCAL) 6 BP 9507 Ouagadougou, Burkina Faso Tel: +226 25375423 Email: secretariat_cc@wascal.org Website: www.wascal.org . Data accessibility Repository name: TOM2024 Data identification number: doi: 10.17632/3d4yg89rtr.1 Direct URL to data: https://data.mendeley.com/datasets/3d4yg89rtr/1 Related research articleOpen asset ↗10.17632/3d4yg89rtr.1lines:1-43
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published3 Feb 2025Frontiers in plant scienceCited by 3 · OpenAlex ↗

Adaptive spatial-channel feature fusion and self-calibrated convolution for early maize seedlings counting in UAV images.

MaizeAerial / UAVField / plotWhole plant / canopy / plot / fieldCountingObject detection

Accurate counting of crop plants is essential for agricultural science, particularly for yield forecasting, field management, and experimental studies. Traditional methods are labor-intensive and prone to errors. Unmanned Aerial Vehicle (UAV) technology offers a promising alternative; however, varying UAV altitudes can impact image quality, leading to blurred features and reduced accuracy in early maize seedling counts. To address these challenges, we developed RC-Dino, a deep learning methodology based on DINO, specifically designed to enhance the precision of seedling counts from UAV-acquired images. RC-Dino introduces two innovative components: a novel self-calibrating convolutional layer named RSCconv and an adaptive spatial feature fusion module called ASCFF. The RSCconv layer improves the representation of early maize seedlings compared to non-seedling elements within feature maps by calibrating spatial domain features. The ASCFF module enhances the discriminability of early maize seedlings by adaptively fusing feature maps extracted from different layers of the backbone network. Additionally, transfer learning was employed to integrate pre-trained weights with RSCconv, facilitating faster convergence and improved accuracy. The efficacy of our approach was validated using the Early Maize Seedlings Dataset (EMSD), comprising 1,233 annotated images of early maize seedlings, totaling 83,404 individual annotations. Testing on this dataset demonstrated that RC-Dino outperformed existing models, including DINO, Faster R-CNN, RetinaNet, YOLOX, and Deformable DETR. Specifically, RC-Dino achieved improvements of 16.29% in Average Precision (AP) and 8.19% in Recall compared to the DINO model. Our method also exhibited superior coefficient of determination (R²) values across different datasets for seedling counting. By integrating RSCconv and ASCFF into other detection frameworks such as Faster R-CNN, RetinaNet, and Deformable DETR, we observed enhanced detection and counting accuracy, further validating the effectiveness of our proposed method. These advancements make RC-Dino particularly suitable for accurate early maize seedling counting in the field. The source code for RSCconv and ASCFF is publicly available at https://github.com/collapser-AI/RC-Dino, promoting further research and practical applications.

Why it matches plant phenotyping methodsUAV画像からトウモロコシ幼苗数を抽出する深層学習手法を開発し、公開データセット上で既存手法と比較検証しており、植物表現型取得法が研究の中心です。

abstractwe developed RC-Dino, a deep learning methodology based on DINO, specifically designed to enhance the precision of seedling counts from UAV-acquired images.
Reproduction assets foundThe paper's EMSD UAV image dataset is explicitly not publicly available, but the authors' RSCconv and ASCFF analysis code for the RC-Dino model is publicly released on GitHub.
Code · publicHowever, our code is open to the public. The RSCconv and ASCFF code mentioned in this paper can be found here: https://github.com/collapser-AI/RC-Dino .Open asset ↗collapser-AI/RC-Dinolines:727-739
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published1 Feb 2025GeneticsCited by 27 · OpenAlex ↗

Global genotype by environment prediction competition reveals that diverse modeling strategies can deliver satisfactory maize yield estimates.

MaizeField / plotYield / biomass estimationYield / yield components

Predicting phenotypes from a combination of genetic and environmental factors is a grand challenge of modern biology. Slight improvements in this area have the potential to save lives, improve food and fuel security, permit better care of the planet, and create other positive outcomes. In 2022 and 2023, the first open-to-the-public Genomes to Fields initiative Genotype by Environment prediction competition was held using a large dataset including genomic variation, phenotype and weather measurements, and field management notes gathered by the project over 9 years. The competition attracted registrants from around the world with representation from academic, government, industry, and nonprofit institutions as well as unaffiliated. These participants came from diverse disciplines, including plant science, animal science, breeding, statistics, computational biology, and others. Some participants had no formal genetics or plant-related training, and some were just beginning their graduate education. The teams applied varied methods and strategies, providing a wealth of modeling knowledge based on a common dataset. The winner's strategy involved 2 models combining machine learning and traditional breeding tools: 1 model emphasized environment using features extracted by random forest, ridge regression, and least squares, and 1 focused on genetics. Other high-performing teams' methods included quantitative genetics, machine learning/deep learning, mechanistic models, and model ensembles. The dataset factors used, such as genetics, weather, and management data, were also diverse, demonstrating that no single model or strategy is far superior to all others within the context of this competition.

Why it matches plant phenotyping methods遺伝・環境情報からトウモロコシ収量という植物形質を予測するモデルを競争形式で比較・評価しており、計算的形質推定とベンチマークが中心である。

titleGlobal genotype by environment prediction competition reveals that diverse modeling strategies can deliver satisfactory maize yield estimates.
Reproduction assets foundThe paper is the G2F maize G×E prediction competition report. Its curated phenotype/genotype/weather/EC dataset is public (DOI 10.25739/tq5e-ak26), but that DOI is not among the allowed URLs, so it cannot be listed. However, the authors explicitly state that code from all participating teams is publicly available, and
Code · publicour abilities to solve critical, and technically challenging, problems. Data availability All data used in this manuscript are publicly available at https:// doi.org/10.25739/tq5e-ak26. Code from all teams is publicly avail­ able as follows: AgAdaptAR: https://github.com/EcoEvoInfo/maize-gxe-pre diction-challenge-2023 AIMaize: https://github.com/ksegaba/Genomes2Field_Competition All Models are Wrong: https://zenodo.org/record/7830071 arulrich: https://github.com/mwylerCH/GxEcompetition CLAC: https://github.com/alenxav/Lectures/tree/master/MGC_2023 DataJanitors: https://github.com/qchen33/g2fcompetition2022 DeepCropVision: https://github.com/Ved-Piyush/DeepCrop Vision_maizegxeprediction2022 EOpen asset ↗ksegaba/Genomes2Field_Competitionpdf-raw-page:13 lines:1-91
Code · publicta availability All data used in this manuscript are publicly available at https:// doi.org/10.25739/tq5e-ak26. Code from all teams is publicly avail­ able as follows: AgAdaptAR: https://github.com/EcoEvoInfo/maize-gxe-pre diction-challenge-2023 AIMaize: https://github.com/ksegaba/Genomes2Field_Competition All Models are Wrong: https://zenodo.org/record/7830071 arulrich: https://github.com/mwylerCH/GxEcompetition CLAC: https://github.com/alenxav/Lectures/tree/master/MGC_2023 DataJanitors: https://github.com/qchen33/g2fcompetition2022 DeepCropVision: https://github.com/Ved-Piyush/DeepCrop Vision_maizegxeprediction2022 EnBiSys: https://github.com/dperondi/maizegxeprediction2022 gartyboiOpen asset ↗pdf-raw-page:13 lines:1-91
Code · publicript are publicly available at https:// doi.org/10.25739/tq5e-ak26. Code from all teams is publicly avail­ able as follows: AgAdaptAR: https://github.com/EcoEvoInfo/maize-gxe-pre diction-challenge-2023 AIMaize: https://github.com/ksegaba/Genomes2Field_Competition All Models are Wrong: https://zenodo.org/record/7830071 arulrich: https://github.com/mwylerCH/GxEcompetition CLAC: https://github.com/alenxav/Lectures/tree/master/MGC_2023 DataJanitors: https://github.com/qchen33/g2fcompetition2022 DeepCropVision: https://github.com/Ved-Piyush/DeepCrop Vision_maizegxeprediction2022 EnBiSys: https://github.com/dperondi/maizegxeprediction2022 gartybois: https://github.com/Thyra/g2f-maize-challenge-202Open asset ↗mwylerCH/GxEcompetitionpdf-raw-page:13 lines:1-91
Code · public0.25739/tq5e-ak26. Code from all teams is publicly avail­ able as follows: AgAdaptAR: https://github.com/EcoEvoInfo/maize-gxe-pre diction-challenge-2023 AIMaize: https://github.com/ksegaba/Genomes2Field_Competition All Models are Wrong: https://zenodo.org/record/7830071 arulrich: https://github.com/mwylerCH/GxEcompetition CLAC: https://github.com/alenxav/Lectures/tree/master/MGC_2023 DataJanitors: https://github.com/qchen33/g2fcompetition2022 DeepCropVision: https://github.com/Ved-Piyush/DeepCrop Vision_maizegxeprediction2022 EnBiSys: https://github.com/dperondi/maizegxeprediction2022 gartybois: https://github.com/Thyra/g2f-maize-challenge-2022 Kernel of Truth: https://github.com/robertkhu/mOpen asset ↗alenxav/Lecturespdf-raw-page:13 lines:1-91
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published31 Jan 2025Data in briefCited by 11 · OpenAlex ↗

A comprehensive image dataset for the identification of eggplant leaf diseases and computer vision applications.

Eggplant / aubergineLeafClassificationDisease symptoms / severity

This dataset on eggplant leaf diseases has been meticulously developed to provide a valuable resource for agricultural research and the advancement of automated disease detection systems. It comprises 4,089 high-resolution images of eggplant leaves, systematically categorized into six distinct classes: Healthy Leaf, Insect Pest Disease, Leaf Spot Disease, Mosaic Virus Disease, White Mold Disease, and Wilt Disease. The images were captured using smartphone cameras under controlled conditions with a consistent white background to ensure clarity and uniformity. To reflect real-world agricultural scenarios, data collection was conducted across multiple geographic locations and in varying lighting conditions. This approach enhances the dataset's diversity and applicability. The dataset underwent thorough manual labelling and preprocessing to ensure accuracy and consistency across all samples. Each image is clearly labelled according to its respective disease class, making the dataset readily usable for machine learning applications. The balanced representation of healthy and diseased leaves allows for comprehensive training and testing of classification models. Designed to support the development of machine learning models for the early detection and classification of eggplant diseases, this dataset holds significant reuse potential in various research domains. It is particularly suitable for applications in plant pathology, precision agriculture, and disease forecasting, where timely and accurate diagnosis is crucial. The dataset is freely available for academic and research purposes, making it a valuable resource for researchers and developers aiming to innovate in agricultural technology and crop management. With its robust design and practical focus, the dataset has the potential to drive advancements in sustainable farming practices and enhance agricultural productivity.

Why it matches plant phenotyping methodsナス葉の病害状態を画像から分類するための大規模データセットであり、植物病害表現型の取得・再利用可能な基盤が中心です。

titleA comprehensive image dataset for the identification of eggplant leaf diseases and computer vision applications.
Reproduction assets foundThe paper is a Data in Brief article describing an eggplant leaf disease image dataset (4,089 images, six classes) publicly deposited on Mendeley Data, plus an authors' GitHub repository containing the preprocessing code. Both are paper-specific, public, and directly actionable.
Dataset · publiclant field in Char Keshabpur, Shibchar, Madaripur (Latitude: 23°21′32.9″N, Longitude: 90°11′48.5″E) 5. Eggplant field in Daffodil Smart City, Khagan, Ashulia (Latitude: 23°52′37.6″N, Longitude: 90°19′16.2″E). Data accessibility Repository name: Mendeley Data. Data identification number: 10.17632/d3ypkphghb.2 Direct URL to data: https://data.mendeley.com/datasets/d3ypkphghb/2 Access the dataset at https://data.mendeley.com/datasets/d3ypkphghb/2 and cite using Data ID 10.17632/d3ypkphghb.2 . Related research articleOpen asset ↗Mendeley Data · 10.17632/d3ypkphghb.2lines:1-49
Code · publicand facilitate classification tasks. • Classification: Images were organized into six predefined classes: Healthy Leaf, Insect Pest, Leaf Spot, Mosaic Virus, White Mold, and Wilt, forming a structured dataset ready for analysis. 4.5. Code used for data preprocessing GitHub Repository name: Data_Preprocessing Direct URL of Code: https://github.com/paradoxicalProfessor/Data_Preprocessing Limitations The Eggplant Leaf Disease dataset has some limitations. It was collected from specific regions in Bangladesh, which may limit its applicability to other environments. Our dataset includes only six disease classes, which may not represent all eggplant diseases in different regions. Some disease clasOpen asset ↗Data_Preprocessinglines:223-258
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published29 Jan 2025Data in briefCited by 6 · OpenAlex ↗

IDBGL: A unique image dataset of black gram (Vigna mungo) leaves for disease detection and classification.

Field / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Black gram (Vigna mungo) is considered one of the most important pulse crops cultivated in Bangladesh because it is a vital source of nutrition and a potential source for raising a good income. It is one of those plants where most leaves are affected by diseases. We observed that most of the leaves were diseased in the fields, and we had difficulty collecting healthy samples. The crop is affected by different diseases attacking leaf tissues, causing heavy yield loss. We can apply deep learning models to recognize diseases in their early stages for timely interference. Diseases could be detected with the automation process, from which much enhancement in the management and yield of black gram crops is possible. Our purpose is to create a unique dataset of Bangladesh's Black Gram (Vigna mungo) to help global researchers build a deep learning-automated system for the early detection and classification of Black Gram leaf diseases that will assist farmers and create more awareness among different agricultural stakeholders. The original dataset of 4,038 images was collected from the Sirajganj and Solonga regions in Bangladesh. The dataset has five different classes: Healthy, Cercospora Leaf Spot, Insect, Leaf Crinkle, and Yellow Mosaic. This dataset will help researchers improve disease detection in Black Grams by developing effective computational models and applying advanced machine learning techniques.

Why it matches plant phenotyping methods黒豆葉の病害状態を画像データセットとして収集・分類する研究で、植物病害フェノタイピング用データセットの構築が中心である。

abstractOur purpose is to create a unique dataset of Bangladesh's Black Gram (Vigna mungo) to help global researchers build a deep learning-automated system for the early detection and classification of Black Gram leaf diseases
Reproduction assets foundThe paper is a data descriptor for a public Mendeley Data image dataset of 4,038 black gram leaf images used for disease detection/classification, with a direct public URL and DOI provided by the authors.
Dataset · publicions in Bangladesh: 1. Black Gram field Sirajganj Sadar, Sirajganj (Latitude: 24°34′40.4"N, Longitude: 89°38′11.7"E) 2. Black Gram field in Solonga, Sirajganj (Latitude: 24°25′08.7"N, Longitude: 89°30′23.5"E). Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/z55yrbmn2d.3 Direct URL to data: https://data.mendeley.com/datasets/z55yrbmn2d/3 1. Value of the Data • This is a dataset dedicated only to black gram (Vigna mungo) leaf diseases, containing valuable resources for both agriculture and machine learning researchers. The classes are well-defined: healthy, Cercospora leaf spot, insect, leaf crinkle, and yellow mosaic. These are diversified sets of imagesOpen asset ↗Mendeley Data · 10.17632/z55yrbmn2d.3lines:1-60
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 Jan 2025Data in briefCited by 4 · OpenAlex ↗

Bean leaf image dataset annotated with leaf dimensions, segmentation masks, and camera calibration.

Common beanLeafMorphology / geometry measurementCalibration / preprocessingSegmentationLeaf traits

Leaf dimensioning is relevant for analyzing plant responses to several conditions such as soil fertility, availability of light, agricultural pesticide effect, and access to water in the soil or periods of drought. In this paper, we present a dataset composed of 6981 images of 612 common bean leaves ( Phaseolus vulgaris ). We captured the images of each leaf accompanied by a fiducial marker and annotated the known leaf dimensions (area, perimeter, length, and width). We provide annotations concerning image segmentation, known area uniformly distributed over the leaf region, real area of the marker region, marker pose, capture conditions, and camera calibration. This dataset can be useful for developing deep learning algorithms for leaf dimensioning and related problems. Therefore, there is a potential to contribute to computer vision and plant physiology researchers and specialists.

Why it matches plant phenotyping methods葉面積・周長・長さ・幅の画像ベース計測用データセットを提供し、セグメンテーション、マーカー姿勢、カメラ校正も含むため、植物表現型取得手法の基盤として中心的です。

abstractWe captured the images of each leaf accompanied by a fiducial marker and annotated the known leaf dimensions (area, perimeter, length, and width).
Reproduction assets foundThe paper is itself a data descriptor for the LSID-Beans bean leaf image dataset (6981 images, 612 leaves, with leaf dimension annotations, segmentation masks, area maps, and camera calibration). The dataset is publicly deposited on Mendeley Data (DOI 10.17632/f42hwwrpgn.2), and the authors' data-processing scripts are
Dataset · publicstakes and improved the data quality. Data source location The images were collected in the city of Ouro Branco, Minas Gerais, Latitude −20.535912, Longitude −43.711031, Brazil. Data accessibility Repository name: Leaf on Stem Image Dataset Beans (LSID-Beans) Data identification number: 10.17632/f42hwwrpgn.2 Direct URL to data: https://data.mendeley.com/datasets/f42hwwrpgn/2 1 Value of the Data • The dataset images are useful for developing deep learning methods for non-destructive leaf dimension estimation. We provide each leaf's known area, perimeter, width, and length, which can be used to train supervised machine learning algorithms. • Methods developed using the dataset can help to moniOpen asset ↗10.17632/f42hwwrpgn.2lines:1-50
Code · publicfor that split. Section Cross-validation protocol definition details our proposed cross-validation protocol. 4 Experimental Design, Materials and Methods Fig. 3 shows the steps performed to build our dataset. We describe each step in the next sections. The source codes used to process the data are available in this repository: https://github.com/gcg-ufjf/LSID-Beans-Scripts . Fig. 3 Steps of the dataset construction. Fig 3 4.1 Plant cultivation We selected black bean seeds and carried out planting in April 2022. On average, 3 seeds were sown in each pit, made with the aid of a hoe, along 9 rows of 30 plants. The soil used had never been cultivated and had rejects of construction material on tOpen asset ↗githublines:66-146
Code / dataset availability confirmedOpenAlex · checked 15 Sept 2026
Published27 Jan 2025Cited by 0 · OpenAlex ↗

Countrywide Digital Surface Models and Vegetation Height Models from Historical Aerial Images

Aerial / UAVPhotogrammetry / SfM / MVSStereo2D/3D reconstructionPlant / canopy height

Abstract. Historical aerial images, captured by film cameras in the previous century, are valuable resources for quantifying Earth’s surface and landscape changes over time. In the post-war period, these images were often acquired to create topographic maps, resulting in the acquisition of large-scale aerial photographs with stereo coverage. Photogrammetric techniques applied to these stereo images enable the extraction of 3D information to reconstruct digital surface models (DSMs) and orthoimages. Here, we present a highly automated photogrammetric approach for generating countrywide DSMs of Switzerland, at a 1 m resolution, from approximately 40,000 scanned aerial stereo images acquired between 1979 and 2006, with known exterior and interior orientation. We derived four countrywide DSMs for the epochs 1979–1985, 1985–1991, 1991–1998, and 1998–2006. From the DSMs, we generated corresponding countrywide vegetation height models (VHMs). We assessed the quality of the historical DSMs at the country scale and within six representative study sites, evaluating the vertical accuracy and the completeness of image-matching across different land cover types. Mean completeness ranged from 64 % for ‘glacial and perpetual snow’ to 98 % for ‘sealed surfaces’, with a value of 93 % for the ‘closed forest’ class. Across Switzerland, the median elevation accuracy of the historical DSMs compared with a reference digital terrain model (DTM) on sealed surface points ranged from 0.28 to 0.53 m, with a normalised median absolute deviation (NMAD) of around 1 m and a maximum root mean square error (RMSE) of 3.90 m. The same analysis between geodetic points and historical DSMs showed higher accuracies, with median values of ≤ 0.05 m and an NMAD < 1 m. The VHMs generated in this study enabled the detection of major changes in forest areas due to windstorm damage, forest dynamics, and growth. This work demonstrates the feasibility of generating accurate, very high-resolution DSM time series (spanning three decades) and VHMs from historical aerial images of the entire surface of Switzerland in a highly automated manner. The VHMs are already being used to estimate countrywide biomass changes. The countrywide DSMs and VHMs for the four epochs, along with auxiliary data, are available online at https://doi.org/10.16904/envidat.528 (Marty et al., 2024) and can be used to quantify long-term elevation changes and related processes across different surfaces.

Why it matches plant phenotyping methods歴史航空画像から植生高モデルを自動生成するフォトグラメトリ手法を開発・精度評価し、森林の高さ変化という植物状態を測定するデータセットも提供しているため、植物フェノタイピング手法が中心である。

abstractFrom the DSMs, we generated corresponding countrywide vegetation height models (VHMs).
Reproduction assets foundThe paper's own countrywide DSM and vegetation height model (VHM) rasters, plus masks and metadata, are publicly deposited on EnviDat with an explicit DOI, directly reproducing the paper's vegetation height measurements.
Dataset · public22 5 Data availability 434 Datasets can be accessed from EnviDat (https://doi.org/10.16904/envidat.528, Marty et al., 2024). The following files are 435 available for the four epochs: countrywide digital surface model (DSM), hillshaded DSM, and vegetation height models 436 (VHMs). A metadata shapefile is provided with information about the acquisition year of the photographs used here; the 437 geometry corresponds to the 1:25,00Open asset ↗EnviDat · 10.16904/envidat.528pdf-raw-page:22 lines:1-63
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published23 Jan 2025Scientific reportsCited by 7 · OpenAlex ↗

A multi-spectral and hyperspectral image dataset for evaluating chemical traits and the water status of avocado, olive and grape through leaf dehydration under laboratory conditions.

AvocadoGrapevineOliveLaboratory / benchtopMultispectral / hyperspectralLeafPhysiological trait estimationPigment / colour / senescenceWater status / transpiration

Assessing the health status of vegetation is of vital importance for all stakeholders. Multi-spectral and hyper-spectral imaging systems are tools for evaluating the health of vegetation in laboratory settings, and also hold the potential of assessing vegetation of large portions of land. However, the literature lacks benchmark datasets to test algorithms for predicting plant health status, with most researchers creating tailored datasets. This work presents a dataset composed of multi-spectral images, hyper-spectral reflectance values, and measurements of weight, chlorophyll, and nitrogen content of leaves at five different drying stages, from avocado, olive, and grape trees, which are common crops in the Valparaíso region of Chile. This dataset is a valuable asset for developing tools in the field of precision agriculture and assessing the general health status of vegetation.

Why it matches plant phenotyping methods植物のマルチスペクトル・ハイパースペクトル画像と葉の水分状態・化学形質を含む評価用データセットを構築しており、フェノタイピング手法開発のためのベンチマークが中心である。

abstractThis work presents a dataset composed of multi-spectral images, hyper-spectral reflectance values, and measurements of weight, chlorophyll, and nitrogen content of leaves at five different drying stages
Reproduction assets foundThe paper's multispectral images, hyperspectral reflectance, and trait measurements (weight, chlorophyll, nitrogen, fuel moisture) are publicly deposited on Figshare with an explicit DOI. The authors' sample Matlab code is included within that dataset. The MicaSense imageprocessing repository is a generic third-party工具
Dataset · publicAll the data is available at this repository DOI: https://doi.org/10.6084/m9.figshare.26950660.v2.Open asset ↗figshare · 10.6084/m9.figshare.26950660.v2pdf-page:14 lines:1-35
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published21 Jan 2025Data in briefCited by 6 · OpenAlex ↗

A comprehensive hog plum leaf disease dataset for enhanced detection and classification.

Field / plotLeafClassificationDisease symptoms / severity

A comprehensive Hog plum leaf disease dataset is greatly needed for agricultural research, precision agriculture, and efficient management of disease. It will find applications toward the formulation of machine learning models for early detection and classification of disease, thus reducing dependency on manual inspections and timely interventions. Such a dataset provides a benchmark for training and testing algorithms, further enhancing automated monitoring systems and decision-support tools in sustainable agriculture. It enables better crop management, less use of chemicals, and more focused agronomical practices. This dataset will contribute to the global research being carried out for the advancement of disease-resistant plant strategy development and efficient management practices for better agricultural productivity along with sustainability. These images have been collected from different regions of Bangladesh. In this work, two classes were used: ' Healthy ' and 'Insect hole' , representing different stages of disease progression. The augmentation techniques that involve flipping, rotating, scaling, translating, cropping, adding noise, adjusting brightness, adjusting contrast, and scaling expanded a dataset of 3782 images to 20,000 images. These have formed very robust deep learning training sets, hence better detection of the disease.

Why it matches plant phenotyping methods植物葉の病徴画像データセットを構築し、疾患検出・分類モデルの訓練とベンチマークに用いることが中心であるため、画像ベースの植物表現型データセットとして採用する。

abstractA comprehensive Hog plum leaf disease dataset is greatly needed for agricultural research, precision agriculture, and efficient management of disease.
Reproduction assets foundThe paper's own hog plum leaf disease image dataset (original and augmented images) is publicly deposited on Mendeley Data with an explicit direct URL and DOI, matching an allowed URL.
Dataset · publicGarden of Bangladesh, Mirpur, Dhaka - 1216. (Latitude: 23° 48′ 46.96″ N, Longitude: 90° 20′ 51.87″ E) 3. Zailla, Singair, Manikganj, Dhaka, Bangladesh. (Latitude: 3° 47′ 46.11″ N, Longitude: 90° 13′ 15.73″ E) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/yvtn2gp8zg .1 Direct URL to data: https://data.mendeley.com/datasets/yvtn2gp8zg/1 Related research article None 1 Value of the Data • The Hog plum leaf disease dataset presents the computer vision and machine learning models with the opportunity to effectively classify leaves between healthy and diseased classes, enabling improved plant health management through early intervention. This dataset helps Open asset ↗Mendeley Datalines:1-56
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published19 Jan 2025BiologyCited by 10 · OpenAlex ↗

Evaluation of Different Few-Shot Learning Methods in the Plant Disease Classification Domain.

ClassificationDisease symptoms / severity

Early detection of plant diseases is crucial for agro-holdings, farmers, and smallholders. Various neural network architectures and training methods have been employed to identify optimal solutions for plant disease classification. However, research applying one-shot or few-shot learning approaches, based on similarity determination, to the plantdisease classification domain remains limited. This study evaluates different loss functions used in similarity learning, including Contrastive, Triplet, Quadruplet, SphereFace, CosFace, and ArcFace, alongside various backbone networks, such as MobileNet, EfficientNet, ConvNeXt, and ResNeXt. Custom datasets of real-life images, comprising over 4000 samples across 68 classes of plant diseases, pests, and their effects, were utilized. The experiments evaluate standard transfer learning approaches alongside similarity learning methods based on two classes of loss function. Results demonstrate the superiority of cosine-based methods over Siamese networks in embedding extraction for disease classification. Effective approaches for model organization and training are determined. Additionally, the impact of data normalization is tested, and the generalization ability of the models is assessed using a special dataset consisting of 400 images of difficult-to-identify plant disease cases.

Why it matches plant phenotyping methods植物病害を画像から分類する少数ショット学習手法を複数比較・評価し、実画像データセットで汎化性能も検証しているため、病害状態のフェノタイピング手法が中心です。

abstractThis study evaluates different loss functions used in similarity learning, including Contrastive, Triplet, Quadruplet, SphereFace, CosFace, and ArcFace, alongside various backbone networks, such as MobileNet, EfficientNet, ConvNeXt, and ResNeXt.
Reproduction assets foundThe paper's reduced-scale (128×128) DoctorP plant disease image dataset is publicly available on Kaggle, as stated in the Data Availability Statement and Dataset section. No author analysis code or trained model checkpoints are reported as publicly available.
Dataset · publicy injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. Institutional Review Board Statement Not applicable. Informed Consent Statement Not applicable. Data Availability Statement The reduced-scale dataset (128 × 128 pixels) is now available for research on Kaggle ( https://www.kaggle.com/datasets/alexanderuzhinskiy/the-doctorp-project-dataset (accessed on 12 December 2024)). Conflicts of Interest The author declares no conflict of interest. References 1. Ramanjot Mittal U. Wadhawan A. Singla J. Jhanjhi N.Z. Ghoniem R.M. Ray S.K. Abdelmaboud A. Plant Disease Detection and Classification: A Systematic Literature Review Sensors 202Open asset ↗Kaggle · the-doctorp-project-datasetlines:107-255
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published17 Jan 2025PloS oneCited by 17 · OpenAlex ↗

Field-scale detection of Bacterial Leaf Blight in rice based on UAV multispectral imaging and deep learning frameworks.

RiceField / plotMultispectral / hyperspectralLeafSegmentationDisease symptoms / severity

Bacterial Leaf Blight (BLB) usually attacks rice in the flowering stage and can cause yield losses of up to 50% in severely infected fields. The resulting yield losses severely impact farmers, necessitating compensation from the regulatory authorities. This study introduces a new pipeline specifically designed for detecting BLB in rice fields using unmanned aerial vehicle (UAV) imagery. Employing the U-Net architecture with a ResNet-101 backbone, we explore three band combinations-multispectral, multispectral+NDVI, and multispectral+NDRE-to achieve superior segmentation accuracy. Due to the lack of suitable UAV-based datasets for rice disease, we generate our own dataset through disease inoculation techniques in experimental paddy fields. The dataset is increased using data augmentation and patch extraction methods to improve training robustness. Our findings demonstrate that the U-Net model incorporating ResNet-101 backbone trained with multispectral+NDVI data significantly outperforms other band combinations, achieving high accuracy metrics, including mean Intersection over Union (mIoU) of up to 97.20%, mean accuracy of up to 99.42%, mean F1-score of up to 98.56%, mean Precision of 97.97%, and mean Recall of 99.16%. Additionally, this approach efficiently segments healthy rice from other classes, minimizing misclassification and improving disease severity assessment. Therefore, the experiment concludes that the accurate mapping of the disease extent and severity level in the field is reliable to accurately allocating the compensation. The developed methodology has the potential for broader application in diagnosing other rice diseases, such as Blast, Bacterial Panicle Blight, and Sheath Blight, and could significantly enhance agricultural management through accurate damage mapping and yield loss estimation.

Why it matches plant phenotyping methodsUAVマルチスペクトル画像と深層学習によるイネ病害の症状・重症度推定パイプラインを開発し、精度評価とデータセット構築を行っており、植物表現型取得が中心である。

abstractThis study introduces a new pipeline specifically designed for detecting BLB in rice fields using unmanned aerial vehicle (UAV) imagery.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicData Availability: We have provided the aerial datasets in figshare after getting permission from the landowner. Please see https://doi.org/10.6084/m9.figshare.26955862.v1 .Open asset ↗figshare · 10.6084/m9.figshare.26955862.v1lines:146-158
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Jan 2025Data in briefCited by 3 · OpenAlex ↗

IDDMSLD: An image dataset for detecting Malabar spinach leaf diseases.

Field / plotLeafClassificationStress / disease detectionDisease symptoms / severity

Agriculture has always played a vital role in the economic development of Bangladesh. In Agriculture, leaf diseases have become an issue because they can lead to a major drop in both quality and quantity of crops. Therefore, leveraging technology to automatically detect diseases on leaves plays an important role in farming. Malabar Spinach (Basella alba) is a well-known, widely grown leafy vegetable, which is valued for its nutritional benefits. However, there is almost no dataset that can aid in identifying diseases affecting this important crop, which often leads to decreased quality as well as financial drawback. This lack of resources makes it difficult for farmers to recognize and manage common diseases. Our purpose is to solve this problem by creating a unique dataset of Bangladesh's Malabar Spinach leaves that will ease agricultural management and disease detection. Our dataset contains both healthy and diseased samples, categorised into four common ailments: Anthracnose, Bacterial Spot, Downy Mildew, and Pest Damage. We collected 3,006 original images in total. Images were collected from various locations in Bangladesh, including Mirpur, Savar, Sirajganj and Gazipur, with photographs taken under natural lighting conditions at different times of the day. This dataset will help the researchers for further research on Malabar Spinach disease detection implementing various efficient computational models and applying advanced machine learning techniques.

Why it matches plant phenotyping methods植物葉の健全・病害状態を画像化したデータセットの作成が研究の中心であり、植物病害フェノタイピング用データセットに該当する。

abstractOur purpose is to solve this problem by creating a unique dataset of Bangladesh's Malabar Spinach leaves that will ease agricultural management and disease detection.
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository of 3,006 Malabar spinach leaf images (healthy plus four disease classes) used for plant disease phenotyping. The dataset is paper-specific, publicly deposited, and directly actionable via the stated direct URL.
Dataset · publicLongitude: 89°30′23.5″E) 3. Malabar Spinach field in Khagan, Ashulia, Savar (Latitude: 23°52′32.1″N, Longitude: 90°19′42.0″E) 4. Malabar Spinach field Gazipur (Latitude: 24°04′14.8″N, Longitude: 90°32′32.3″E) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/sy69db2nz5.2 Direct URL to data: https://data.mendeley.com/datasets/sy69db2nz5/2 1. Value of the Data • This dataset valuable because it provides a large, diverse collection of images of Malabar Spinach (Basella alba) affected by common diseases. This diverse collection will help in disease classification and modeling specifically for this crop as well as for the agricultural research and better diseaOpen asset ↗Mendeley Data · 10.17632/sy69db2nz5.2lines:1-55
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 Jan 2025Data in briefCited by 2 · OpenAlex ↗

Annotated image dataset with different stages of European pear rust for UAV-based automated symptom detection in orchards.

PearAerial / UAVField / plotRGB / grayscaleLeafObject detectionDisease symptoms / severity

The evaluation of fruit genetic resources regarding a resistance to pathogens is an essential basis for subsequent selection in fruit breeding. Both genetic analysis and phenotyping of defined traits are important tools and provide decision data in the evaluation process. However, the phenotyping of plants is often carried out 'by hand' and remains the bottleneck in fruit breeding and fruit growing. The development of a digital and UAV (unmanned aerial vehicle)-based phenotyping method for the assessment of genotype-specific susceptibility or resistance against diseases in orchards would significantly increase the efficiency of plant breeding. In this framework, a workflow for drone-based monitoring of pathogens in orchards was developed using the European pear rust ( Gymnosporangium sabinae ) as model pathogen. Pear rust is widespread in orchards and causes conspicuous, clearly visible, yellow to orange-colored disease symptoms. In this paper, we provide a dataset with expert-annotated high-resolution RGB images with pear rust symptoms. For data collection, ten UAV-flight campaigns were realized between 2021 and 2023 under various weather conditions and with different flight parameters in the experimental orchard of the Julius Kühn-Institute for Breeding Research on Fruit Crops in Dresden-Pillnitz (Germany). 1394 images were captured of different pear genotypes, including varieties, wild species and progeny from breeding. The dataset contains manually labelled images with a size of 768 × 768 pixels of leaves infected with pear rust at different stages of development, labelled as class GYMNSA, as well as background images without symptoms. Each leaf with pear rust symptoms was annotated with the drawing method by two points (bounding boxes) using the Computer Vision Annotation Tool (CVAT, v1.1.0) [1] and presented in YOLO 1.1 file format (.txt files). A total of 584 annotated images and 162 background images, organized into a training and validation set, are included in the GYMNSA dataset. This GYMNSA dataset can be used as a resource for researchers and developers working on drone-based plant disease monitoring systems.

Why it matches plant phenotyping methodsナシさび病の植物症状をUAV画像から検出するための注釈付きデータセットを提供しており、植物病害状態の画像ベース表現型取得・解析ワークフローが中心的です。

abstractThe development of a digital and UAV (unmanned aerial vehicle)-based phenotyping method for the assessment of genotype-specific susceptibility or resistance against diseases in orchards would significantly increase the efficiency of plant breeding.
Reproduction assets foundThe paper's GYMNSA dataset — annotated UAV RGB images of pear rust symptoms with YOLO labels — is publicly deposited on Mendeley Data under DOI 10.17632/44kjgc4gkc.1, directly reproducing the paper's phenotyping measurements.
Dataset · publicl orchard of the Julius Kühn-Institute (JKI - Federal Research Centre for Cultivated Plants) at the Institute for Breeding Research on Fruit Crops located in Dresden-Pillnitz (Germany) [51°00ʹ01"N 13°53ʹ12"E]. Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/44kjgc4gkc.1 Direct URL to data: https://data.mendeley.com/datasets/44kjgc4gkc/1 1. Value of the Data • These data were collected on an approximately 1.6 ha experimental field with over 1000 different pear genotypes (breeding material and genetic resources of pear varieties and species) and presents a wide spectrum of phenotypic characteristics of pear rust infections at different stages of developmeOpen asset ↗Mendeley Data · 10.17632/44kjgc4gkc.1lines:43-69
Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Published1 Jan 2025DatabaseCited by 0 · OpenAlex ↗

CPDMS: a database system for crop physiological disorder management

TomatoWhole plant / canopy / plot / fieldObject detectionDisease symptoms / severityStress response / tolerance

Abstract As the importance of precision agriculture grows, scalable and efficient methods for real-time data collection and analysis have become essential. In this study, we developed a system to collect real-time crop images, focusing on physiological disorders in tomatoes. This system systematically collects crop images and related data, with the potential to evolve into a valuable tool for researchers and agricultural practitioners. A total of 58 479 images were produced under stress conditions, including bacterial wilt (BW), Tomato Yellow Leaf Curl Virus (TYLCV), Tomato Spotted Wilt Virus (TSWV), drought, and salinity, across seven tomato varieties. The images include front views at 0 degrees, 120 degrees, 240 degrees, and top views and petiole images. Of these, 43 894 images were suitable for labeling. Based on this, 24 000 images were used for AI model training, and 13 037 images for model testing. By training a deep learning model, we achieved a mean Average Precision (mAP) of 0.46 and a recall rate of 0.60. Additionally, we discussed data augmentation and hyperparameter tuning strategies to improve AI model performance and explored the potential for generalizing the system across various agricultural environments. The database constructed in this study will serve as a crucial resource for the future development of agricultural AI. Database URL: https://crops.phyzen.com/

Why it matches plant phenotyping methodsトマトの生理障害・病害を対象に画像収集データベースと深層学習解析モデルを開発しており、植物状態の取得・推定手法が研究の中心である。

titleCPDMS: a database system for crop physiological disorder management
Reproduction assets foundThe paper's tomato physiological-disorder image dataset (58,479 images, annotations, and AI training data) is publicly available via the authors' CPDMS database. LabelImg and YOLOv5 are generic third-party tools, not paper-specific assets.
Dataset · publicAll data used in this study are publicly available at https://crops.phyzen.com/ and https://crops.phyzen.com/appOpen asset ↗crops.phyzen.comlines:141-251
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 15 Sept 2026
Published1 Jan 2025GigaScienceCited by 28 · OpenAlex ↗

High-fidelity wheat plant reconstruction using 3D Gaussian splatting and neural radiance fields

WheatField / plotNeRF / 3D Gaussian SplattingLiDAR / point cloudRGB / grayscaleWhole plant / canopy / plot / fieldCalibration / preprocessing2D/3D reconstruction

BACKGROUND: The reconstruction of 3-dimensional (3D) plant models can offer advantages over traditional 2-dimensional approaches by more accurately capturing the complex structure and characteristics of different crops. Conventional 3D reconstruction techniques often produce sparse or noisy representations of plants using software or are expensive to capture in hardware. Recently, view synthesis models have been developed that can generate detailed 3D scenes, and even 3D models, from only RGB images and camera poses. These models offer unparalleled accuracy but are currently data hungry, requiring large numbers of views with very accurate camera calibration. RESULTS: In this study, we present a view synthesis dataset comprising 20 individual wheat plants captured across 6 different time frames over a 15-week growth period. We develop a camera capture system using 2 robotic arms combined with a turntable, controlled by a re-deployable and flexible image capture framework. We trained each plant instance using two recent view synthesis models: 3D Gaussian splatting (3DGS) and neural radiance fields (NeRF). Our results show that both 3DGS and NeRF produce high-fidelity reconstructed images of a plant subject from views not captured in the initial training sets. We also show that these approaches can be used to generate accurate 3D representations of these plants as point clouds, with 0.74-mm and 1.43-mm average accuracy compared with a handheld scanner for 3DGS and NeRF, respectively. CONCLUSION: We believe that these new methods will be transformative in the field of 3D plant phenotyping, plant reconstruction, and active vision. To further this cause, we release all robot configuration and control software, alongside our extensive multiview dataset. We also release all scripts necessary to train both 3DGS and NeRF, all trained models data, and final 3D point cloud representations. Our dataset can be accessed via https://plantimages.nottingham.ac.uk/ or https://https://doi.org/10.5524/102661. Our software can be accessed via https://github.com/Lewis-Stuart-11/3D-Plant-View-Synthesis.

Why it matches plant phenotyping methods3D植物表現型取得のための撮影システム、再構成手法、データセットを開発し、スキャナとの精度比較で検証しているため、方法が中心的である。

abstractWe develop a camera capture system using 2 robotic arms combined with a turntable, controlled by a re-deployable and flexible image capture framework.
Reproduction assets foundThe paper releases its wheat plant multiview image dataset (via plantimages.nottingham.ac.uk and GigaDB DOI 10.5524/102661), its authors' analysis/capture codebase on GitHub (3D-Plant-View-Synthesis), a Software Heritage archive of that code, and a DOME-ML registry annotation. All are paper-specific, public, and have作者
Code · publicruction output across all plants. We hope that our study will provide opportunities for researchers exploring new and improved 3D phenotyping algorithms, 3D reconstruction and view synthesis research, and active vision systems. Availability of Source Code and Requirements Project name: 3D Plant View Synthesis: Project homepage: https://github.com/Lewis-Stuart-11/3D-Plant-View-Synthesis [ 13 ] Operating system(s): Windows, Ubuntu Programming language: Python (>=3.8) License: Apache 2.0 Any restrictions to use by nonacademics: None Our code has also been archived in Software Heritage [ 66 ]. Functionality, such as Robotic View Capturing, 3DGS to Point Cloud, and our UR5 Configs files, are storOpen asset ↗GitHub · Lewis-Stuart-11/3D-Plant-View-Synthesislines:663-695
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published27 Dec 2024Scientific dataCited by 27 · OpenAlex ↗

Crops3D: a diverse 3D crop dataset for realistic perception and segmentation toward agricultural applications.

LiDAR / point cloudWhole plant / canopy / plot / fieldClassificationSegmentation

Point cloud analysis is a crucial task in computer vision. Despite significant advances over the past decade, the developments in agricultural domain have faced challenges due to a scarcity of datasets. To facilitate 3D point cloud research in agriculture community, we introduce Crops3D, the diverse real-world dataset derived from authentic agricultural scenarios. Crops3D distinguishes itself through its unique properties: diversity, authenticity, and complexity. The dataset incorporates data from diverse point cloud acquisition methods, encompassing eight distinct crop types with 1,230 samples, authentically representing crops in the real-world. It stands as the pioneering dataset that comprehensively supports the three critical tasks in 3D crop phenotyping: instance segmentation of individual plants in agricultural settings, plant type perception, and plant organ segmentation. Additionally, the intricate crop structures in Crops3D exhibit higher complexity than available 3D public datasets, showcasing substantial self-occlusion and increased complexity as crops mature. We analyse diverse crop point cloud acquisition methods and evaluate multiple models' performance with the Crops3D dataset.

Why it matches plant phenotyping methods3D作物点群データセットを構築し、個体・器官のセグメンテーションなど植物フェノタイピング用途で複数モデルと取得法を評価しており、データセットと解析手法が中心である。

abstractIt stands as the pioneering dataset that comprehensively supports the three critical tasks in 3D crop phenotyping: instance segmentation of individual plants in agricultural settings, plant type perception, and plant organ segmentation.
Reproduction assets foundThe paper's Crops3D point cloud dataset is deposited in figshare, but no figshare URL is among the allowed URLs, so the dataset itself cannot be linked. The authors' analysis code (subsampling, corruption, S3DIS-format conversion scripts, environment files) is explicitly stated to be publicly available on GitHub, which
Code · publicthe subsampling, corruption scripts, conversion to S3DIS format scripts, along with other code-related content, are available through the following GitHub repository: https://github.com/clawCa/Crops3DOpen asset ↗https://github.com/clawCa/Crops3Dpdf-page:15 lines:1-21
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published27 Dec 2024Data in briefCited by 7 · OpenAlex ↗

Smartphone image dataset for radish plant leaf disease classification from Bangladesh.

RadishField / plotRGB / grayscaleLeafClassificationDisease symptoms / severity

Radishes, which are common root vegetables, are rich in vitamins and minerals, and contain low calories. This vegetable is known for its rapid growth. Nevertheless, the variety of leaf diseases where leaves get affected by various bacterial and fungal diseases can hinder the healthy growth of radish. Furthermore, there is a high risk of inaccurate identification of diseases if the farmers try to use traditional methods in recognizing these diseases. With the purpose of precise identification of radish leaf diseases for the finest growth of this vegetable, total of 2801 images of the radish leaves are collected from vegetable field in Bangladesh. The collected dataset includes comprehensive images of healthy leaves as well as four types of leaf affected by various diseases such as Black Leaf Spot, Downey Mildew, Flea Beetle and Mosaic. Utilizing this robust dataset, deep learning models can be trained to identify the leaf diseases which helps to detect the diseases in order to reduce the harm of the cultivation of radish. By identifying the diseases on radish leaves accurat-ely and maintaining healthy production of radish, this dataset contributes to the broader sustainability in the agricultural sector.

Why it matches plant phenotyping methodsダイコン葉の病害状態を画像で取得したデータセットの構築が中心であり、植物病害フェノタイピング用の再利用可能な資源に該当する。

abstracttotal of 2801 images of the radish leaves are collected from vegetable field in Bangladesh
Reproduction assets foundThe paper is a Data in Brief article describing a public Mendeley Data repository of 2801 smartphone images of radish leaves (healthy plus four disease classes) collected in Bangladesh, which is the paper's own phenotyping image dataset and is directly accessible.
Dataset · publicortant role for classifying the radish plant healthy and unhealthy leaves. Data source location 1. Vegetable field of Kathalkandi, Nasirnagar, Brahmanbaria, Bangladesh (latitude: 24.1915°, longitude: 91.1826°) Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/s973cz2jcd.1 Direct URL to data: https://data.mendeley.com/datasets/s973cz2jcd/1 1 Value of the Data • The dataset containing several classes of radish leaves where each class clearly representing the unhealthy leaf as well as healthy leaf. All the images are captured with high resolution that ensuing the high-quality of leaves images, helps to recognize the pattens of diseases. • The dataset presentOpen asset ↗Mendeley Data · 10.17632/s973cz2jcd.1lines:1-50
Code / dataset availability confirmedEurope PMC · checked 13 Sept 2026
Published24 Dec 2024Data in briefCited by 4 · OpenAlex ↗

Dataset of aerial photographs acquired with UAV using a multispectral (green, red and near-infrared) camera for cherry tomato ( Solanum lycopersicum var. cerasiforme ) monitoring.

CherryTomatoAerial / UAVField / plotMultispectral / hyperspectralWhole plant / canopy / plot / fieldClassificationObject detection2D/3D reconstructionSegmentation

A dataset of aerial photographs acquired with an Unmanned Aerial Vehicle (UAV) DJI Phantom 4 Pro is presented for monitoring a cherry tomato ( Solanum lycopersicum var. cerasiforme ) crop in Navolato, Mexico. Seven photogrammetric flights were carried out to assess the plant growth using a Mapir Survey 3W multispectral camera. Multispectral images with an approximate spatial resolution of 1.83 cm/px were obtained in each photogrammetric flight. These images were acquired every 15 days starting on October 15, 2021, and ending on January 23, 2022. The dataset contains the radiometrically calibrated images of the tomato crop divided into 2 open field parcels. The dataset also includes the processed photogrammetric products (ortho-mosaics) using a binary mask to exclude the soil from the plant area. The dataset was originally acquired to assess plant growth, stress levels, and overall crop health. However, this multispectral imagery dataset can also have various uses, such as creating training datasets with accurate labels or classes which can then be used to develop, train, and/or validate machine learning algorithms for image classification, object detection tasks, or change detection analysis.

Why it matches plant phenotyping methods植物の生育・ストレス・健全性評価を目的とした、放射補正済みマルチスペクトル画像とオルソモザイクを含む再利用可能なデータセットであり、植物表現型取得基盤が中心です。

abstractThe dataset contains the radiometrically calibrated images of the tomato crop divided into 2 open field parcels.
Reproduction assets foundThe paper is itself a data descriptor for a public UAV multispectral cherry tomato phenotyping dataset (calibrated aerial images, manual plant images, orthomosaics, binary masks) deposited in Dryad, with an explicit DOI and direct URL matching an allowed URL.
Dataset · publicRepository name: tomatodb Data identification number: 10.5061/dryad.63xsj3vbd Direct URL to data: https://datadryad.org/stash/share/Wq_X7QUyGryJ-ZnmgfwRn4MtOCr4VBm_MSnhF40sv_8#readmeOpen asset ↗Dryad · 10.5061/dryad.63xsj3vbdlines:1-42
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 6 Sept 2026
Published19 Dec 2024Plant PhenomicsCited by 24 · OpenAlex ↗

From Images to Loci: Applying 3D Deep Learning to Enable Multivariate and Multitemporal Digital Phenotyping and Mapping the Genetics Underlying Nitrogen Use Efficiency in Wheat

WheatAerial / UAVField / plotLiDAR / point cloudMultispectral / hyperspectralWhole plant / canopy / plot / fieldMorphology / geometry measurementSegmentationGrowth / time-series analysisGrowth / development / phenology

The selection and promotion of high-yielding and nitrogen-efficient wheat varieties can reduce nitrogen fertilizer application while ensuring wheat yield and quality and contribute to the sustainable development of agriculture; thus, the mining and localization of nitrogen use efficiency (NUE) genes is particularly important, but the localization of NUE genes requires a large amount of phenotypic data support. In view of this, we propose the use of low-altitude aerial photography to acquire field images at a large scale, generate 3-dimensional (3D) point clouds and multispectral images of wheat plots, propose a wheat 3D plot segmentation dataset, quantify the plot canopy height via combination with PointNet++, and generate 4 nitrogen utilization-related vegetation indices via index calculations. Six height-related and 24 vegetation-index-related dynamic digital phenotypes were extracted from the digital phenotypes collected at different time points and fitted to generate dynamic curves. We applied height-derived dynamic numerical phenotypes to genome-wide association studies of 160 wheat cultivars (660,000 single-nucleotide polymorphisms) and found that we were able to locate reliable loci associated with height and NUE, some of which were consistent with published studies. Finally, dynamic phenotypes derived from plant indices can also be applied to genome-wide association studies and ultimately locate NUE- and growth-related loci. In conclusion, we believe that our work demonstrates valuable advances in 3D digital dynamic phenotyping for locating genes for NUE in wheat and provides breeders with accurate phenotypic data for the selection and breeding of nitrogen-efficient wheat varieties.

Why it matches plant phenotyping methods航空画像・3D点群・マルチスペクトル画像から小麦区画の草冠高と植生指数を抽出するデジタルフェノタイピング手法を開発・適用しており、表現型取得が研究の中心である。

abstractwe propose the use of low-altitude aerial photography to acquire field images at a large scale, generate 3-dimensional (3D) point clouds and multispectral images of wheat plots
Reproduction assets foundThe paper's Data Availability statement explicitly deposits the authors' source code, testing data, and supporting datasets (including the W3DPS 3D plot segmentation dataset and phenotyping/GWAS data) at two public Quark pan links under CC BY 4.0. These are paper-specific, publicly actionable assets. Other allowed URLs
Code · publicThe source code, testing data, and other datasets supporting the results presented here are available at https://pan.quark.cn/s/afbf9025b19e and https://pan.quark.cn/s/47e91f9d6c9c .Open asset ↗lines:138-156
Dataset · publicThe source code, testing data, and other datasets supporting the results presented here are available at https://pan.quark.cn/s/afbf9025b19e and https://pan.quark.cn/s/47e91f9d6c9c .Open asset ↗lines:138-156
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published19 Dec 2024Data in briefCited by 25 · OpenAlex ↗

A comprehensive image dataset for the identification of lemon leaf diseases and computer vision applications.

CitrusField / plotLeafClassificationDisease symptoms / severity

A comprehensive dataset on lemon leaf disease can surely bring a lot of potentials into the development of agricultural research and the improvement of disease management strategies. This dataset was developed from 1354 raw images taken with professional agricultural specialist guidance from July to September 2024 in Charpolisha, Jamalpur, and further enhanced with augmented techniques, adding 9000 images. The augmentation process involves a set of techniques-flipping, rotation, zooming, shifting, adding noise, shearing, and brightening-to increase variety for different lemon leaf condition representations. Each of these images was standardized to 800 × 800 pixels resolution, so that consistency may be maintained among the dataset. All images were labelled in the nine prefixed categories: anthracnose, bacterial blight, citrus canker, curl virus, deficiency leaf, dry leaf, healthy leaf, sooty mould, and spider mites. In the present study, a DenseNet-121 architecture was used, where 20 % of the dataset was kept for validation and the remaining 80 % for training. A trained model with a batch size of 32 was trained for 30 epochs, achieving an accuracy of 98.56 % with augmentation, and 96.19 % without it. The dataset will not only act as a benchmark in developing accurate machine learning models for early disease detection, but it will also contribute to the cause of sustainable lemon cultivation practices by facilitating timely and effective disease management interventions .

Why it matches plant phenotyping methodsレモン葉の病害・健全状態を画像で表現するデータセットを構築し、分類性能を検証しており、植物病害表現型の取得・ベンチマークが中心です。

abstractA comprehensive dataset on lemon leaf disease can surely bring a lot of potentials into the development of agricultural research and the improvement of disease management strategies.
Reproduction assets foundThe paper's own lemon leaf disease image dataset (1354 original + 9000 augmented images) is publicly deposited on Mendeley Data with DOI 10.17632/44nrn4593f.1 and a direct URL, making it a paper-specific, publicly actionable asset.
Dataset · publicder mites. Since then, the collection of images has been highly varied, which is good enough for deep learning applications. Data source location Town/City/Region: Charpolisha, Jamalpur. Country: Bangladesh . Data accessibility Repository name: Mendeley Data. Data identification number: 10.17632/44nrn4593f.1 Direct URL to data: https://data.mendeley.com/datasets/44nrn4593f/1 Related research article None . 1 Value of the Data • The dataset contains various images of lemon leaves infected with different diseases, right from the most common to the rare ones. Thus, it will be very helpful in agriculture and scientific aspects for extending research in plant pathology. This dataset thus finds itOpen asset ↗Mendeley Data · 10.17632/44nrn4593f.1lines:1-49
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published16 Dec 2024Plant phenomics (Washington, D.C.)Cited by 8 · OpenAlex ↗

Informed-Learning-Guided Visual Question Answering Model of Crop Disease.

ClassificationStress / disease detectionDisease symptoms / severity

In contemporary agriculture, experts develop preventative and remedial strategies for various disease stages in diverse crops. Decision-making regarding the stages of disease occurrence exceeds the capabilities of single-image tasks, such as image classification and object detection. Consequently, research now focuses on training visual question answering (VQA) models. However, existing studies concentrate on identifying disease species rather than formulating questions that encompass crucial multiattributes. Additionally, model performance is susceptible to the model structure and dataset biases. To address these challenges, we construct the informed-learning-guided VQA model of crop disease (ILCD). ILCD improves model performance by integrating coattention, a multimodal fusion model (MUTAN), and a bias-balancing (BiBa) strategy. To facilitate the investigation of various visual attributes of crop diseases and the determination of disease occurrence stages, we construct a new VQA dataset called the Crop Disease Multi-attribute VQA with Prior Knowledge (CDwPK-VQA). This dataset contains comprehensive information on various visual attributes such as shape, size, status, and color. We expand the dataset by integrating prior knowledge into CDwPK-VQA to address performance challenges. Comparative experiments are conducted by ILCD on the VQA-v2, VQA-CP v2, and CDwPK-VQA datasets, achieving accuracies of 68.90%, 49.75%, and 86.06%, respectively. Ablation experiments are conducted on CDwPK-VQA to evaluate the effectiveness of various modules, including coattention, MUTAN, and BiBa. These experiments demonstrate that ILCD exhibits the highest level of accuracy, performance, and value in the field of agriculture. The source codes can be accessed at https://github.com/SdustZYP/ILCD-master/tree/main.

Why it matches plant phenotyping methods作物病害の視覚属性と発生段階を画像から推定するVQAモデルと専用データセットを開発しており、植物状態の表現型推定手法が研究の中心である。

abstractwe construct the informed-learning-guided VQA model of crop disease (ILCD).
Reproduction assets foundThe paper's authors publicly release both the ILCD analysis code and the paper-specific CDwPK-VQA dataset (crop disease images with question–answer annotations) via GitHub URLs stated in the article.
Code · publicd 86.06%, respectively. Ablation experiments are conducted on CDwPK-VQA to evaluate the effectiveness of various modules, including coattention, MUTAN, and BiBa. These experiments demonstrate that ILCD exhibits the highest level of accuracy, performance, and value in the field of agriculture. The source codes can be accessed at https://github.com/SdustZYP/ILCD-master/tree/main. status released display-pdf yes is-olf no is-manuscript no is-preprint no is-journal-matter no is-scanned no is-retracted no Received 2024 May 24; Revised 2024 Oct 18; Accepted 2024 Nov 12; Collection date 2024. Introduction The Food and Agriculture Organization of the United Nations has reports that diseases are respOpen asset ↗SdustZYP/ILCD-masterlines:1-26
Dataset · publicgnment between the question text information and image region features. This process results prior knowledge dataset comprising 272 images and 2,180 questions. CDwPK-VQA integrates prior knowledge to expand the dataset and regulate the learning behavior of the model, as shown in Fig. 2 . The dataset of CDwPK-VQA is available at https://github.com/SdustZYP/CDwPK-VQA/tree/main. The ILCD model This research constructs a novel ILCD. The model architecture of ILCD is shown in Fig. 3 , and divided into the following steps: (a) Image features V and question features Q are extracted using a pretrained feature extraction model. (b) The coattention mechanism captures the interaction between the image Open asset ↗SdustZYP/CDwPK-VQAlines:52-91
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published12 Dec 2024Data in briefCited by 3 · OpenAlex ↗

Lentil plant disease and quality assessment: A detailed dataset of high-resolution images for deep learning research.

LentilField / plotWhole plant / canopy / plot / fieldStress / disease detectionDisease symptoms / severity

The Lentil, a vital legume globally cultivated, faces significant challenges from diseases like ascochyta blight, lentil rust, and powdery mildew. Ensuring optimal harvest timing and effectively discerning healthy and diseased lentil plants are crucial for maintaining crop quality and economic viability, particularly in regions such as Bangladesh. This paper introduces a comprehensive dataset comprising high-resolution images of lentil plants gathered meticulously over four months from diverse locations across Bangladesh, under expert supervision. The dataset aims to support the development of machine-learning models for precise disease detection and quality assessment in lentil cultivation. Potential applications include enhancing the accuracy of quality evaluation, and improving packaging processes, thereby enhancing overall lentil production efficiency. Agricultural researchers can utilize this dataset to advance applications of computer vision and deep learning in managing crop diseases and enhancing yield outcomes. The dataset's creation involved collaboration with domain experts to ensure its relevance and reliability for agricultural research. By leveraging this dataset, researchers can explore innovative approaches to tackle challenges in lentil farming, contributing to sustainable agricultural practices and food security. Moreover, the dataset serves as a valuable resource for training and testing machine learning algorithms tailored to agricultural settings, facilitating advancements in automated agricultural technologies. Ultimately, this initiative aims to empower stakeholders in the lentil industry with tools to mitigate disease impact and optimize production practices, paving the way for more resilient and efficient agricultural systems globally .

Why it matches plant phenotyping methodsレンティル植物の病害状態を画像から評価する高解像度データセットを構築し、機械学習による病害検出を支援することが中心であり、植物フェノタイピング用データセットに該当する。

abstractThis paper introduces a comprehensive dataset comprising high-resolution images of lentil plants
Reproduction assets foundThe paper is a Data in Brief article describing a lentil plant disease image dataset (1,898 original and 4,550 augmented images across four classes) deposited publicly on Mendeley Data with DOI 10.17632/7vb77bz2st.1. This is a paper-specific, publicly accessible image dataset directly reproducing the paper's phenotypic
Dataset · publicnd research stations in Barisal, Bangladesh Coordinates: 1. Farm A: 22.7083° N, 90.3653° E 2. Farm B: 22.6715° N, 90.3252° E 3. Research Station C: 22.7032° N, 90.3863° E Zone: Barisal] Country: [Bangladesh] Data accessibility Repository name: [Mendeley Data] Data identification number: 10.17632/7vb77bz2st.1 Direct URL to data: https://data.mendeley.com/datasets/7vb77bz2st/1 Related research article [None] How the dataset helps in packaging 1. [ Quality Assessment : Automates the assessment of lentil quality, ensuring only high-quality products are packaged. 2. Sorting and Grading : Aids in developing algorithms for sorting lentils based on size, color, and disease presence, enhancing efficiOpen asset ↗Mendeley Data · 10.17632/7vb77bz2st.1lines:1-55
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Dec 2024Data in briefCited by 4 · OpenAlex ↗

Money plant disease atlas: A comprehensive dataset for disease classification in ornamental horticulture.

LeafClassificationStress / disease detectionDisease symptoms / severity

Epipremnum aureum, sometimes known as the Money Plant, is a popular houseplant known for its hearts-shaped leaves and durability. Commonly referred to as Golden Pothos or Devil's Ivy, it is also appreciated for its ornamental value and air cleaning ability. They say that these plants are attractive to many people owing to their tolerance to several conditions and easy care, therefore, it is no surprise that they are found in many households and workplaces. Money Plants are hardy, but like any other plant they can also be infected by various diseases, which may render them less attractive, or even unattractive. This work encompasses bacterial wilt, manganese poisoning aspects and together with a healthy leaves aspect presents all prevalent masses and offer a comprehensive image of diseases. A dataset of 224 × 224 pixel images is utilized to accomplish this work with the intention to further enhance support in Ornamental Horticulture practices and diagnose more accurately. This work not only contributes ideas and approaches in understanding the field of plants pathology but also stresses on the fact how image processing can be beneficial in looking after plants. The dataset serves as a solid foundation for deep learning approaches into Ornamental Agriculture and provides useful insights for researchers studying the cultivation of money plants.

Why it matches plant phenotyping methods植物病害の症状を画像から分類するデータセットが研究の中心であり、罹病状態という植物表現型を直接評価している。

abstractA dataset of 224 × 224 pixel images is utilized to accomplish this work with the intention to further enhance support in Ornamental Horticulture practices and diagnose more accurately.
Reproduction assets foundThe paper's money plant leaf disease image dataset (original and augmented) is publicly deposited on Mendeley Data, and the authors' data augmentation code is publicly available on GitHub; both are paper-specific, public, and directly actionable.
Dataset · publicen in collaboration with an expert from the Ministry of Agriculture, Bangladesh Data source location Location: Bangladesh Agricultural Development Corporation. The area: Kashimpur, Gazipur. Country: Bangladesh Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/rzjww3vdxt.3 Direct URL to data: https://data.mendeley.com/datasets/rzjww3vdxt/3 Related research article None 1. Value of the Data Epipremnum aureum, popularly called Devil's ivy, Golden images, or Money Plant, is a common houseplant that is appreciated for its durability as well as its heart shaped leaves [ 1 ]. Known not only for its pleasing aesthetic but also for its air purifying qualities, theOpen asset ↗Mendeley Data · 10.17632/rzjww3vdxt.3lines:1-46
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published6 Dec 2024Earth System Science DataCited by 6 · OpenAlex ↗

Observational partitioning of water and CO 2 fluxes at National Ecological Observatory Network (NEON) sites: a 5-year dataset of soil and plant components for spatial and temporal analysis

Field / plotWhole plant / canopy / plot / fieldGrowth / time-series analysisPhotosynthesis / fluorescenceWater status / transpiration

Abstract. Long-term time series of transpiration, evaporation, plant net photosynthesis, and soil respiration are essential for addressing numerous research questions related to ecosystem functioning. However, quantifying these fluxes is challenging due to the lack of reliable and direct measurement techniques, which has left gaps in the understanding of their temporal cycles and spatial variability. To help address this open challenge, we generated a dataset of these four components by implementing five (conventional and novel) approaches to partition total evapotranspiration (ET) and CO2 fluxes into plant and soil fluxes across 47 National Ecological Observatory Network (NEON) sites. The final dataset (https://doi.org/10.5281/zenodo.12191876; Zahn and Bou-Zeid, 2024) spans a 5-year period and covers various ecosystems, including forests, grasslands, and agricultural terrain. This is the first comprehensive dataset covering such a wide spatial and temporal distribution. Overall, we observed good agreement across most methods for ET components, increasing confidence in these estimates. Partitioning of CO2 components, on the other hand, was found to be less robust and more dependent on prior knowledge of water use efficiency. This highlights some limitations of these present methods that we discuss, emphasizing the broader challenge posed by the lack of an accurate reference method to validate against. Despite these limitations, this dataset has several potential applications, especially in addressing critical questions regarding the response of ecosystems to extreme weather events, which are expected to become more severe and frequent with climate change.

Why it matches plant phenotyping methods複数手法で蒸散・植物純光合成などの植物生理フラックスを分離推定し、手法間比較と妥当性・限界評価を行った大規模データセットであり、植物状態の取得手法が中心である。

abstractwe generated a dataset of these four components by implementing five (conventional and novel) approaches to partition total evapotranspiration (ET) and CO2 fluxes into plant and soil fluxes across 47 National Ecological Observatory Network (NEON) sites.
Reproduction assets foundThe paper's flux-partitioning dataset (transpiration, evaporation, plant photosynthesis, soil respiration across 47 NEON sites) is publicly deposited on Zenodo, and the authors' scripts implementing all five partitioning methods are also publicly available on Zenodo with explicit availability statements.
Dataset · publicl. ( 2020 ) across FLUXNET sites. By comparing different algorithms, we can further explore their uncertainties and focus on model improvement. Finally, as more data become available, other options can be used to train machine learning algorithms, focusing on gap-filling. 7 Code and data availability The dataset is available at https://doi.org/10.5281/zenodo.12191876 ( Zahn and Bou-Zeid , 2024 ) . In addition to all the flux components, it contains the auxiliary meteorological inputs used to implement the Extreme Gradient Boosting algorithm for gap-filling and feature importance analysis. The scripts used to implement all five partitioning methods can be found at https://doi.org/10.5281/zenOpen asset ↗Zenodo · 10.5281/zenodo.12191876lines:346-361
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published5 Dec 2024Data in briefCited by 1 · OpenAlex ↗

Image dataset: UAV images and ground data of one 'Bingo' mandarin and two 'Valencia' orange rootstock trials conducted in Florida.

CitrusAerial / UAVField / plotFruitWhole plant / canopy / plot / fieldArchitecture / morphology / geometryPlant / canopy heightYield / yield components

The data are aerial images and ground tree measurement data of 3 citrus rootstock trials. Developing new citrus rootstock varieties requires field trials to test to identify selections with improved horticultural performance. A bud from a scion variety is grafted onto the rootstock and grown in a nursery until the grafted plant is ready to be planted in the field, which is in about one year. Trees in the field are assessed each year by measuring height, canopy diameter in 2 dimensions, overall health, and fruit number and quality factors when the trees begin to have a significant crop (∼3 years). Data collection of each tree is done manually. The image and ground data sets are of 3 rootstock trials that includes a 3-year-old Bingo mandarin hybrid trial of 206 trees, a 6-year-old Valencia orange trial of 643 trees, and a 7-year-old Valencia orange trials of 648 trees. Data for each trial includes aerial images and ground data of height, canopy diameters, and an overall health rating. The combination of ground validated measures and aerial images make this data set useful for building AI-based aerial image data collection applications. The data will be useful for 1) visualizing the effects of different rootstock selections and varieties on scion growth, effects that may not be fully captured with single measure metrics; and 2) development of image analysis applications and segmentation algorithms that can extract data from the images that are suitable for replacing some or all the ground measures.

Why it matches plant phenotyping methods柑橘樹の高さ、樹冠径、健康状態を対象とする航空画像・地上測定データセットで、画像解析やセグメンテーションによる形質抽出の開発用途が明示されており、表現型取得法が中心である。

abstractThe combination of ground validated measures and aerial images make this data set useful for building AI-based aerial image data collection applications.
Reproduction assets foundThis Data in Brief article describes its own paper-specific phenotyping assets: UAV RGB images and ground-measured canopy height/width/health data for three citrus rootstock trials, publicly deposited in USDA Ag Data Commons under DOIs 10.15482/USDA.ADC/26946823 (Bingo trial) and 10.15482/USDA.ADC/26946841 (Valencia 5–
Dataset · publicRepository name: USDA Ag Data Commons [ 1 ] Direct URL to Rows 1–4 Bingo rootstock data: 10.15482/USDA.ADC/26946823USDA Ag Data Commons · 10.15482/USDA.ADC/26946823lines:1-53
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published4 Dec 2024Cited by 0 · OpenAlex ↗

A multi-spectral and hyperspectral image dataset for evaluating the health status of avocado, olive and vineyard

AvocadoGrapevineOliveMultispectral / hyperspectralLeafStress / disease detectionPigment / colour / senescenceWater status / transpiration

Abstract Assessing the health status of vegetation is of vital importance for all stakeholders. Multi-spectral and hyper-spectral imaging systems are tools for evaluating the health of crops across large areas, particularly when deployed on robotic platforms such as unmanned aerial vehicles (UAVs). However, the literature lacks benchmark datasets to test algorithms for predicting plant health status, with most researchers creating tailored datasets. This work presents a dataset composed of multi-spectral images, hyper-spectral reflectance values, and measurements of weight, chlorophyll, and nitrogen content of leaves at five different drying stages, from avocado, olive, and vineyard trees, which are common crops in the Valparaíso region of Chile. This dataset is a valuable asset for developing tools in the field of precision agriculture and assessing the general health status of vegetation.

Why it matches plant phenotyping methods植物の健康状態を推定するためのマルチスペクトル・ハイパースペクトル画像と葉の形質測定を組み合わせた評価用データセットが主題であり、植物フェノタイピング手法のベンチマーク資源に該当する。

abstractThis work presents a dataset composed of multi-spectral images, hyper-spectral reflectance values, and measurements of weight, chlorophyll, and nitrogen content of leaves at five different drying stages
Reproduction assets foundThe paper is a dataset descriptor; its complete plant-phenotyping measurements (multispectral leaf images, hyperspectral reflectance, chlorophyll, nitrogen, weight/FMC across five drying stages for avocado, olive, and vineyard) are publicly deposited on figshare under DOI 10.6084/M9.FIGSHARE.26950660, along with aMatlå
Dataset · publicAll the data is available at this repository DOI: 10.6084/M9.FIGSHARE.26950660Open asset ↗figshare · 10.6084/M9.FIGSHARE.26950660pdf-page:15 lines:1-59
Code / dataset availability confirmedCrossref · Europe PMC · checked 13 Sept 2026
Published1 Dec 2024Data in BriefCited by 4 · OpenAlex ↗

Comprehensive smartphone image dataset for bean and cowpea plant leaf disease detection and freshness assessment from Bangladesh vegetable fields

Common beanCowpeaField / plotLeafClassificationObject detectionStress / disease detectionDisease symptoms / severity

Agriculture greatly impacts Bangladesh's economy, and vegetable cultivation plays a significant role in Agriculture by providing nourishment, and food security as well as improving the economy. The necessity of food production is growing similarly to the population growth. The farmers of Bangladesh are working hard to meet this need for food production and to gain yields. However, every year the farmers face a significant amount of loss in production due to the attack of different diseases and viruses due to the lack to technological development. The reason behind most of these losses is the lack of knowledge about diseases and being unable to detect the diseases early. Therefore, the early detection of plant disease is significant in balancing the country's economy and preventing undesirable losses. To bring a solution to this problem our dataset provides a total of 4467 images of Beans and Cowpeas leaf images which include different disease classes and fresh leaves. The dataset comprises 2,273 images of Bean and 2,194 images of Cowpea plants where each plant provides 4 classes of different disease along with the healthy leaves. This dataset will assist researchers in identifying plant diseases and farmers as well as contribute to the economy of the country.

Why it matches plant phenotyping methods豆類葉の画像から病害状態を推定する画像データセットが研究の中心であり、植物病害表現型のデータ資源として収録対象です。

titleComprehensive smartphone image dataset for bean and cowpea plant leaf disease detection and freshness assessment from Bangladesh vegetable fields
Reproduction assets foundThe paper is a Data in Brief article describing a smartphone image dataset of bean and cowpea leaf disease/freshness. The authors' own dataset is publicly deposited on Mendeley Data with an explicit direct URL and DOI, making it a paper-specific, publicly actionable asset. The Kaggle bean disease dataset is cited prior
Dataset · publicData accessibility Repository name: Mendeley Data Data identification number: 10.17632/ykvcrjffzd.1 Direct URL to data: https://data.mendeley.com/datasets/ykvcrjffzd/1Open asset ↗Mendeley Data · 10.17632/ykvcrjffzd.1lines:1-51
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 7 Sept 2026
Published26 Nov 2024Plant PhenomicsCited by 20 · OpenAlex ↗

PlanText: Gradually Masked Guidance to Align Image Phenotypes with Trait Descriptions for Plant Disease Texts

Stress / disease detectionDisease symptoms / severity

Plant diseases are a critical driver of the global food crisis. The integration of advanced artificial intelligence technologies can substantially enhance plant disease diagnostics. However, current methods for early and complex detection remain challenging. Employing multimodal technologies, akin to medical artificial intelligence diagnostics that combine diverse data types, may offer a more effective solution. Presently, the reliance on single-modal data predominates in plant disease research, which limits the scope for early and detailed diagnosis. Consequently, developing text modality generation techniques is essential for overcoming the limitations in plant disease recognition. To this end, we propose a method for aligning plant phenotypes with trait descriptions, which diagnoses text by progressively masking disease images. First, for training and validation, we annotate 5,728 disease phenotype images with expert diagnostic text and provide annotated text and trait labels for 210,000 disease images. Then, we propose a PhenoTrait text description model, which consists of global and heterogeneous feature encoders as well as switching-attention decoders, for accurate context-aware output. Next, to generate a more phenotypically appropriate description, we adopt 3 stages of embedding image features into semantic structures, which generate characterizations that preserve trait features. Finally, our experimental results show that our model outperforms several frontier models in multiple trait descriptions, including the larger models GPT-4 and GPT-4o. Our code and dataset are available at https://plantext.samlab.cn/.

Why it matches plant phenotyping methods植物病害画像から表現型・形質記述を生成するモデルと注釈付きデータセットを開発し、性能比較まで行っており、植物表現型の取得・抽出が研究の中心である。

abstractwe propose a method for aligning plant phenotypes with trait descriptions, which diagnoses text by progressively masking disease images.
Reproduction assets foundThe paper's plant disease image–text dataset (17,183 annotated images, 103,098 labels, 5,728 expert texts) and the PhenoTrait/PlanText code are publicly released at the authors' site https://plantext.samlab.cn, and the authors' data annotation platform code is public on GitHub at https://github.com/kej-shas/data-angles
Code · publicwe develop a data annotation platform (the platform has built-in functions such as image display, text translation, and exporting of Word files; the code is at https://github.com/kej-shas/data-annotations )Open asset ↗kej-shas/data-annotationslines:77-85
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published23 Nov 2024Data in briefCited by 2 · OpenAlex ↗

A comprehensive dataset of near infrared spectroscopy measurements to predict nitrogen and carbon contents in a wide range of tissues from Brassica napus plants grown under contrasted environments.

Rapeseed / canolaRaman / spectroscopyTissuePhysiological trait estimation

Winter oilseed rape (WOSR, Brassica napus L.) is the third largest oil crop worldwide that also provides a source of high quality plant-based proteins. Nitrogen (N) and carbon (C) play a key role in plant growth. Determination of N and C contents of plant tissues throughout the growth cycle is crucial in assessing plant nutritional status and allowing precise input management. In the dataset presented in this article, 2427 WOSR samples arising from a large diversity of tissues collected on WOSR diversity were analyzed by near infrared spectroscopy from 4000 to 12,000 cm -1 . At the same time, reference chemical data for the N and C contents of the same samples were determined by elemental analysis using the Dumas method. Partial least squares regression has been used to develop predictive models linking spectral and chemical data, so that new samples can be characterized without the need for reference methods. This dataset could be used to test new calculation algorithms in order to enhance prediction performance or for training purposes. These models can be used as a rapid method for determining N and/or C content, adding to decision-support tools for fertilizer application throughout the plant developmental cycle.

Why it matches plant phenotyping methods植物組織の窒素・炭素含量を近赤外分光で推定する予測モデルと大規模データセットが研究の中心であり、植物形質・栄養状態の取得手法として実質的です。

abstractIn the dataset presented in this article, 2427 WOSR samples arising from a large diversity of tissues collected on WOSR diversity were analyzed by near infrared spectroscopy
Reproduction assets foundThe article is a Data in Brief describing a paper-specific public dataset of NIR spectra and N/C reference measurements for 2427 Brassica napus tissue samples, deposited in Data INRAE with an explicit DOI and direct URL. The dataset includes the raw spectral data (.csv), chemical reference data, and the PLS calibration
Dataset · publicData source location Institution: Institute of Genetics, Environment and Plant Protection (IGEPP); INRAE, Institut Agro, University of Rennes City/Town/Region: 35,650 Le Rheu Country: France Data accessibility Repository name: Data INRAE ( https://data.inrae.fr/ ) Data identification number: 10.57745/6VYUQN Direct URL to data: https://entrepot.recherche.data.gouv.fr/dataset.xhtml?persistentId=doi:10.57745/6VYUQN Related research article None 1 Value of the Data • The dataset establishes a link between spectral properties and chemical composition (N, C) of a wide variety of plant tissues in winter oilseed rape. The prediction models can be used by diverse communities (scientists, breeders, prOpen asset ↗Data INRAE · 10.57745/6VYUQNlines:1-63
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published20 Nov 2024BMC plant biologyCited by 4 · OpenAlex ↗

NIRSpredict: a platform for predicting plant traits from near infra-red spectroscopy.

ArabidopsisRaman / spectroscopyPhysiological trait estimation

Near-infrared spectroscopy (NIRS) has become a popular tool for investigating phenotypic variability in plants. We developed the Shiny NIRSpredict application to get predictions of 81 Arabidopsis thaliana phenotypic traits, including classical functional traits as well as a large variety of commonly measured chemical compounds, based from near-infrared spectroscopy values based on deep learning. It is freely accessible at the following URL: https://shiny.cefe.cnrs.fr/NirsPredict/ . NIRSpredict has three main functionalities. First, it allows users to submit their spectrum values to get the predictions of plant traits from models built with the hosted A. thaliana database. Second, users have access to the database of traits used for model calibration. Data can be filtered and extracted on user's choice and visualized in a global context. Third, a user can submit his own dataset to extend the database and get part of the application development. NIRSpredict provides an easy-to-use and efficient method for trait prediction and an access to a large dataset of A. thaliana trait values. In addition to covering many of functional traits it also allows to predict a large variety of commonly measured chemical compounds. As a reliable way of characterizing plant populations across geographical ranges, NIRSpredict can facilitate the adoption of phenomics in functional and evolutionary ecology.

Why it matches plant phenotyping methodsNIRスペクトルから植物形質を予測するソフトウェアおよびデータベースを開発しており、形質取得・推定手法が研究の中心である。

abstractWe developed the Shiny NIRSpredict application to get predictions of 81 Arabidopsis thaliana phenotypic traits
Reproduction assets foundThe paper's NIRS spectra and 81 trait measurements for 5,325 Arabidopsis thaliana individuals are publicly hosted in the authors' NIRSpredict Shiny application, and the application's R code is deposited on the authors' GitHub repository (AxelVaillant/NirsPredict), as stated in the Data availability section. Both are直接,
Dataset · publicTrait values are publicly available in the NIRSpredict database atOpen asset ↗pdf-page:10 lines:1-65
Code · publicThe R code of the application is available on a GitHubOpen asset ↗pdf-page:10 lines:1-65
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published15 Nov 2024Scientific reportsCited by 17 · OpenAlex ↗

Integrating deep learning for visual question answering in Agricultural Disease Diagnostics: Case Study of Wheat Rust.

WheatLeafClassificationStress / disease detectionDisease symptoms / severity

This paper presents a novel approach to agricultural disease diagnostics through the integration of Deep Learning (DL) techniques with Visual Question Answering (VQA) systems, specifically targeting the detection of wheat rust. Wheat rust is a pervasive and destructive disease that significantly impacts wheat production worldwide. Traditional diagnostic methods often require expert knowledge and time-consuming processes, making rapid and accurate detection challenging. We drafted a new, WheatRustDL2024 dataset (7998 images of healthy and infected leaves) specifically designed for VQA in the context of wheat rust detection and utilized it to retrieve the initial weights on the federated learning server. This dataset comprises high-resolution images of wheat plants, annotated with detailed questions and answers pertaining to the presence, type, and severity of rust infections. Our dataset also contains images collected from various sources and successfully highlights a wide range of conditions (different lighting, obstructions in the image, etc.) in which a wheat image may be taken, therefore making a generalized universally applicable model. The trained model was federated using Flower. Following extensive analysis, the chosen central model was ResNet. Our fine-tuned ResNet achieved an accuracy of 97.69% on the existing data. We also implemented the BLIP (Bootstrapping Language-Image Pre-training) methods that enable the model to understand complex visual and textual inputs, thereby improving the accuracy and relevance of the generated answers. The dual attention mechanism, combined with BLIP techniques, allows the model to simultaneously focus on relevant image regions and pertinent parts of the questions. We also created a custom dataset (WheatRustVQA) with our augmented dataset containing 1800 augmented images and their associated question-answer pairs. The model fetches an answer with an average BLEU score of 0.6235 on our testing partition of the dataset. This federated model is lightweight and can be seamlessly integrated into mobile phones, drones, etc. without any hardware requirement. Our results indicate that integrating deep learning with VQA for agricultural disease diagnostics not only accelerates the detection process but also reduces dependency on human experts, making it a valuable tool for farmers and agricultural professionals. This approach holds promise for broader applications in plant pathology and precision agriculture and can consequently address food security issues.

Why it matches plant phenotyping methods小麦葉の画像からさび病の有無・種類・重症度を推定するVQA、データセット、連合学習モデルを開発・評価しており、植物病害状態の取得が中心的な方法貢献である。

abstractThis dataset comprises high-resolution images of wheat plants, annotated with detailed questions and answers pertaining to the presence, type, and severity of rust infections.
Reproduction assets foundThe paper's data availability statement explicitly releases the authors' FL/VQA code on GitHub and the paper-specific wheat rust image datasets (WheatRustDL2024, WheatRustVQA images and question-answer text) via public SharePoint/Google Drive/Docs links.
Dataset · public• This study introduces a Federated Learning and a Visual Question-Answering model. These models are available online on this study’s GitHub (https://github.com/aknnvt/FL-VQA-in-Wheat-Rust). • The custom datasets curated for this study, WheatRustDL2024 (https://bitspilaniac-my.sharepoint.com/:f:/g/personal/f20212378_pilani_bits-pilani_ac_in/EvwjsY_JT4FIu7ZTU8zyXOMB8Ywk4OXgO6LYwTk8dOiN_Q? e=45eTOD) and WheatRustVQA (https://docs.google.com/document/d/1EvVdrMi7W-JZkeeePEkNmVIEJ1dn7eln/edit? usp=sharing&ouid=114090611032812705334&rtpof=true&sd=true), are available for public use. Additionally, the images in WheatRustVQA (https://drive.google.com/drive/folders/1izs5ZVmi9V__ixk4ODiJAyachAulP3RL? Open asset ↗WheatRustDL2024lines:292-351
Dataset · publicmodels are available online on this study’s GitHub (https://github.com/aknnvt/FL-VQA-in-Wheat-Rust). • The custom datasets curated for this study, WheatRustDL2024 (https://bitspilaniac-my.sharepoint.com/:f:/g/personal/f20212378_pilani_bits-pilani_ac_in/EvwjsY_JT4FIu7ZTU8zyXOMB8Ywk4OXgO6LYwTk8dOiN_Q? e=45eTOD) and WheatRustVQA (https://docs.google.com/document/d/1EvVdrMi7W-JZkeeePEkNmVIEJ1dn7eln/edit? usp=sharing&ouid=114090611032812705334&rtpof=true&sd=true), are available for public use. Additionally, the images in WheatRustVQA (https://drive.google.com/drive/folders/1izs5ZVmi9V__ixk4ODiJAyachAulP3RL? usp=drive_link). The raw version of the answers can be found on the GitHub repository. DecOpen asset ↗WheatRustVQAlines:292-351
Dataset · public/g/personal/f20212378_pilani_bits-pilani_ac_in/EvwjsY_JT4FIu7ZTU8zyXOMB8Ywk4OXgO6LYwTk8dOiN_Q? e=45eTOD) and WheatRustVQA (https://docs.google.com/document/d/1EvVdrMi7W-JZkeeePEkNmVIEJ1dn7eln/edit? usp=sharing&ouid=114090611032812705334&rtpof=true&sd=true), are available for public use. Additionally, the images in WheatRustVQA (https://drive.google.com/drive/folders/1izs5ZVmi9V__ixk4ODiJAyachAulP3RL? usp=drive_link). The raw version of the answers can be found on the GitHub repository. Declarations Competing interests The authors declare no competing interests. References 1. Abebe W Wheat Leaf Rust Disease Management: a review J. Plant. Pathol. Microbiol. 2021 12 1 8 Abebe, W. Wheat Leaf RusOpen asset ↗WheatRustVQAlines:292-351
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published14 Nov 2024Data in briefCited by 1 · OpenAlex ↗

Microscopy and transcriptomic datasets for investigating the drought-stress response and recovery in young and early senescent-old leaves from Brassica napus .

Rapeseed / canolaMicroscopyCell / cellular structureLeafTissueSegmentationStress / disease detectionLeaf traitsStress response / tolerance

The present dataset combines transcriptomic and microscopic analyses to investigate the responses of winter oilseed rape (WOSR, Brassica napus L., cultivar Aviso) to soil drought, with a focus on differences between young and early-senescent old leaves. For microscopy, 36 scans of 1 to 5 leaf cross-sections were acquired from paraffin-embedded leaf disc samples using a scanner with a 40x lens (Pannoramic Confocal, 3DHistech), capturing a large field of view (8-mm-long observed leaf tissue). The raw scanned cross-sections and analyzed images are available under doi.org/10.57745/RK5PM3 in the Recherche Data Gouvrepository. These high-quality scans enable the differentiation of mesophyll cells and tissues. Software analysis yielded a dataset with 54 selected cross-sectional areas, 291 delimited surfaces of palisade, spongy, and vessel tissues, and 11,136 individually delimited cells from the palisade and spongy layers. For transcriptomics, an Illumina Novaseq sequencer was used to generate 390 Gb of mRNA paired-end reads. The raw reads were filtered, mapped, and assigned to genes from the Brassica napus reference genome Darmor-bzh v10, which were subsequently used to identify differentially expressed genes (DEGs) and to perform gene ontology enrichment analysis. The raw reads are accessible under accession PRJNA939927 at the NCBI Sequence Read Archive (SRA). This high-quality dataset provides insights into the molecular mechanisms underlying oilseed rape's response to soil drought and may aid in the development of drought-tolerant cultivars. A total of 17,975 DEGs were identified between well-watered and severe drought conditions across the contrasted leaf developmental stages.

Why it matches plant phenotyping methods葉の断面画像を取得・解析し、組織面積や個別細胞などの植物形態形質を構造化した再利用可能なデータセットを提供しており、画像ベースの表現型取得が実質的な構成要素である。

abstractFor microscopy, 36 scans of 1 to 5 leaf cross-sections were acquired from paraffin-embedded leaf disc samples using a scanner with a 40x lens
Reproduction assets foundThe article deposits its own plant-phenotyping assets publicly: raw and analyzed leaf cross-section microscopy scans (Recherche Data Gouv, doi:10.57745/RK5PM3) and the transcriptomic dataset (Recherche Data Gouv doi:10.57745/7HQSM3, mirrored at NCBI SRA under PRJNA939927). The analysis pipelines cited (nf-core/rnaseq,
Dataset · publicThe raw scanned cross-sections and analyzed images are available under doi.org/10.57745/RK5PM3 in the Recherche Data Gouvrepository.Open asset ↗Recherche Data Gouv · 10.57745/RK5PM3lines:1-41
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published8 Nov 2024Plant phenomics (Washington, D.C.)Cited by 5 · OpenAlex ↗

Counting Canola: Toward Generalizable Aerial Plant Detection Models.

Rapeseed / canolaAerial / UAVWhole plant / canopy / plot / fieldCountingObject detection

Plant population counts are highly valued by crop producers as important early-season indicators of field health. Traditionally, emergence rate estimates have been acquired through manual counting, an approach that is labor-intensive and relies heavily on sampling techniques. By applying deep learning-based object detection models to aerial field imagery, accurate plant population counts can be obtained for much larger areas of a field. Unfortunately, current detection models often perform poorly when they are faced with image conditions that do not closely resemble the data found in their training sets. In this paper, we explore how specific facets of a plant detector's training set can affect its ability to generalize to unseen image sets. In particular, we examine how a plant detection model's generalizability is influenced by the size, diversity, and quality of its training data. Our experiments show that the gap between in-distribution and out-of-distribution performance cannot be closed by merely increasing the size of a model's training set. We also demonstrate the importance of training set diversity in producing generalizable models, and show how different types of annotation noise can elicit different model behaviors in out-of-distribution test sets. We conduct our investigations with a large and diverse dataset of canola field imagery that we assembled over several years. We also present a new web tool, Canola Counter, which is specifically designed for remote-sensed aerial plant detection tasks. We use the Canola Counter tool to prepare our annotated canola seedling dataset and conduct our experiments. Both our dataset and web tool are publicly available.

Why it matches plant phenotyping methods航空画像からカノーラ個体数(個体群密度)を推定する検出モデルの汎化性能を検証し、注釈付きデータセットと専用Webツールを提示しており、植物表現型取得手法が中心である。

abstractBy applying deep learning-based object detection models to aerial field imagery, accurate plant population counts can be obtained for much larger areas of a field.
Reproduction assets foundThe paper's aerial canola seedling dataset (images and annotations) is publicly deposited on Zenodo, and the authors' Canola Counter analysis/annotation tool is open source on GitHub. The arXiv 2108.05789 entry is a cited prior work (CropAndWeed dataset), not a paper-specific asset.
Dataset · publicData Availability Statement The canola seedling dataset used in this study is publicly available and can be found at: https://doi.org/10.5281/zenodo.11055599 . The Canola Counter tool is open source and is available at: https://github.com/eandvaag/agricounter .Open asset ↗zenodo · 10.5281/zenodo.11055599lines:236-237
Code · publicData Availability Statement The canola seedling dataset used in this study is publicly available and can be found at: https://doi.org/10.5281/zenodo.11055599 . The Canola Counter tool is open source and is available at: https://github.com/eandvaag/agricounter .Open asset ↗github · eandvaag/agricounterlines:236-237
Code / dataset availability confirmedOpenAlex · checked 14 Sept 2026
Published3 Nov 2024The Plant Phenome JournalCited by 4 · OpenAlex ↗

Manifold and spatiotemporal learning on multispectral unoccupied aerial system imagery for phenotype prediction

RiceMultispectral / hyperspectralWhole plant / canopy / plot / fieldGrowth / time-series analysisYield / biomass estimationGrowth / development / phenologyYield / yield components

Abstract Timeseries data captured by unoccupied aircraft systems (UASs) are increasingly used for agricultural applications requiring accurate prediction of plant phenotypes from remotely sensed imagery. However, prediction models often fail to generalize well from one year to the next or to new environments. Here, we investigate the ability of various machine learning (ML) approaches to improve yield prediction accuracy in new environments from multispectral timeseries imagery acquired on a set of rice (Oryza sativa L.) experiments with different management treatments and varieties. We also trained deep learning models that perform automated feature extraction and compared these against a suite of other approaches. We observed similar performance on a held‐out growing season for a spatiotemporal model (a three‐dimensional convolutional neural network) trained on raw images compared to simpler workflows using dimension reduction of manually extracted features from temporal imagery (i.e., vegetation indices and image texture properties). Manifold learning on raw imagery was better suited for the prediction of phenological traits due to the preservation of local structure in image embeddings at some time points. Together, these results highlight the competitiveness of classical ML approaches for UAS image analysis alongside computationally expensive deep learning models. Along with a new benchmark dataset for rice, our results help extend the toolkit for UAS image analysis, contributing to improved phenotype prediction in plant breeding and precision agriculture applications.

Why it matches plant phenotyping methodsUASマルチスペクトル時系列画像から収量・生育期形質を予測する機械学習手法を比較・評価し、米のベンチマークデータセットも提供しており、表現型取得・推定手法が研究の中心である。

abstractprediction models often fail to generalize well from one year to the next or to new environments.
Reproduction assets foundThe paper's data availability statement explicitly deposits raw and processed UAS imagery, extracted features, and agronomic data on Dryad, and the authors' analysis code on GitHub. Both are paper-specific, public, and actionable.
Dataset · publicts complied with the current laws of the United States, the country in which they were performed. C O N F L I C T O F I N T E R E S T S TAT E M E N T Emily S. Bellis is a full time employee of Avalo, Inc., a crop improvement company. DATA AVA I L A B I L I T Y S TAT E M E N T Raw and processed UAS images are available on Dryad (https://doi.org/10.5061/dryad.v41ns1s4z) along with extracted features and agronomic data for the 2021 and 2022 field seasons. Code to reproduce the analyses are available at https://github.com/FareedFarag/TPPJ-Modeling-Code.O RC I D FaredFarag https://orcid.org/0000-0002-4659-6781 Trevis D. Huggins https://orcid.org/0000-0002-1937-6687 JeremyD. Edwards https://orcidOpen asset ↗Dryad · 10.5061/dryad.v41ns1s4zpdf-raw-page:16 lines:1-86
Code · public61/dryad.v41ns1s4z) along with approaches for rice trait prediction using UAS imagery as extracted features and agronomic data for the 2021 and 2022 the primary data source. While showcasing the potential of field seasons. Code to reproduce the analyses are available at various modeling approaches, it also emphasizes the trade- https://github.com/FareedFarag/TPPJ-Modeling-Code. offs between performance and interpretability for applications in precision agriculture and plant breeding. Looking for- ORCID ward, extending the study over multiple years, extending to Fared Farag https://orcid.org/0000-0002-4659-6781 hyperspectral sensors, and exploring additional remotely Trevis D. Huggins https:/Open asset ↗GitHub · FareedFarag/TPPJ-Modeling-Codepdf-layout-page:16 lines:1-54
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 Nov 2024Data in briefCited by 1 · OpenAlex ↗

High spatial and spectral resolution dataset of hyperspectral look-up tables for 3.5 million traits and structural combinations of Central European temperate broadleaf forests.

Field / plotMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldCalibration / preprocessingArchitecture / morphology / geometryLeaf traitsPigment / colour / senescence

Accurate retrieval of forest functional traits from remote sensing data is critical for monitoring forest health and productivity. To achieve sufficient accuracy using inverse methods it is essential to have representative database of simulated or measured spectral properties together with corresponding forest traits. However, existing datasets are often limited in scope, covering specific sites and times with simplified structures. This limitation hinders the development of generalizable machine learning models for trait prediction. To address this issue, we present a comprehensive high-resolution dataset of hyperspectral Look-Up Tables (LUT) designed for Central European temperate broadleaf forests. The dataset includes 3.5 million unique combinations of leaf biochemical and canopy structural characteristics of forest scenes together with a variety of sun geometry. The spectral data cover wavelengths from 450 nm to 2300 nm, with a resolution of 2 nm. The dataset is organised into two files: one capturing the average reflectance of all scene pixels and another focusing solely on sunlit leaf pixels. LUT were generated using the Discrete Anisotropic Radiative Transfer model version 5.10.0. Virtual forest scenes were based on 3D tree representations derived from Terrestrial Laser Scanning of European beech trees, adjusted to various leaf area index values and structural configurations to simulate natural forest variability. The reflectance data were processed using MATLAB and Python scripts, resulting in hyperspectral cubes that were processed to generate the LUT. The dataset can be used to train machine learning models, such as Random Forest and Support Vector Machines, for predicting forest functional traits and assisting in the calibration of remote sensing algorithms. The biggest advantage of the dataset is high spectral and spatial resolution, together with the high number of different trait combinations, which allows for adaptability to different times, locations, and hyper- and multispectral sensors, and can support up-coming hyperspectral satellite missions. ESA Copernicus Hyperspectral Imaging Mission for the Environment (CHIME) and NASA Surface Biology and Geology (SBG) future satellite missions can utilise this dataset to develop their product processors for monitoring forest traits.

Why it matches plant phenotyping methods森林の機能形質を推定するための大規模ハイパースペクトルLUTデータセットを構築しており、形質取得・推定基盤そのものが研究の中心である。

abstractwe present a comprehensive high-resolution dataset of hyperspectral Look-Up Tables (LUT) designed for Central European temperate broadleaf forests.
Reproduction assets foundThe paper is a Data in Brief article describing a public hyperspectral LUT dataset (3.5 million trait/structural combinations) deposited in the Czech National Repository, including the authors' processing codes (merge_images.m, LUT_processing.py) within the deposit. The direct repository URL is given in the text and is
Dataset · publiceaf pixels. Data source location Institutions: Institute of Computer Science, Masaryk University; Global Change Research Institute of the Czech Academy of Sciences City: Brno Country: Czech Republic Data accessibility Repository name: National Repository Data identification number: 10.48700/datst.bcnpf-47q73 Direct URL to data: https://data.narodni-repozitar.cz/general/datasets/4y0sy-qh735 1. Value of the Data • Look-Up Tables (LUT) are considered important training datasets for machine learning models to predict leaf traits. • To date, only a limited number of LUT datasets have been developed for forest sites, particularly for Central European temperate broadleaf forests. Most of them are lOpen asset ↗National Repository · 10.48700/datst.bcnpf-47q73lines:36-69
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published22 Oct 2024Sensors (Basel, Switzerland)Cited by 7 · OpenAlex ↗

Enhancing Grapevine Node Detection to Support Pruning Automation: Leveraging State-of-the-Art YOLO Detection Models for 2D Image Analysis.

GrapevineField / plotStem / branchObject detectionArchitecture / morphology / geometry

Automating pruning tasks entails overcoming several challenges, encompassing not only robotic manipulation but also environment perception and detection. To achieve efficient pruning, robotic systems must accurately identify the correct cutting points. A possible method to define these points is to choose the cutting location based on the number of nodes present on the targeted cane. For this purpose, in grapevine pruning, it is required to correctly identify the nodes present on the primary canes of the grapevines. In this paper, a novel method of node detection in grapevines is proposed with four distinct state-of-the-art versions of the YOLO detection model: YOLOv7, YOLOv8, YOLOv9 and YOLOv10. These models were trained on a public dataset with images containing artificial backgrounds and afterwards validated on different cultivars of grapevines from two distinct Portuguese viticulture regions with cluttered backgrounds. This allowed us to evaluate the robustness of the algorithms on the detection of nodes in diverse environments, compare the performance of the YOLO models used, as well as create a publicly available dataset of grapevines obtained in Portuguese vineyards for node detection. Overall, all used models were capable of achieving correct node detection in images of grapevines from the three distinct datasets. Considering the trade-off between accuracy and inference speed, the YOLOv7 model demonstrated to be the most robust in detecting nodes in 2D images of grapevines, achieving F1-Score values between 70% and 86.5% with inference times of around 89 ms for an input size of 1280 × 1280 px. Considering these results, this work contributes with an efficient approach for real-time node detection for further implementation on an autonomous robotic pruning system.

Why it matches plant phenotyping methodsブドウの節という明示的な植物器官形質をYOLO画像解析で検出する手法を開発・比較検証し、異なる品種・環境で評価しているため、フェノタイピング手法が中心である。

abstractIn this paper, a novel method of node detection in grapevines is proposed with four distinct state-of-the-art versions of the YOLO detection model: YOLOv7, YOLOv8, YOLOv9 and YOLOv10.
Reproduction assets foundThe paper's authors created and openly released a paper-specific grapevine node-detection image dataset (Dão and Douro vineyard images) on Zenodo, cited in the Data Availability Statement.
Dataset · publicThe data presented in this study are openly available in the digital repository Zenodo: Douro & Dão Grapevines Dataset for Node Detection— https://doi.org/10.5281/zenodo.10991688 .Open asset ↗Zenodo · 10.5281/zenodo.10991688lines:552-565
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published18 Oct 2024Scientific reportsCited by 9 · OpenAlex ↗

Deep learning based approach for actinidia flower detection and gender assessment.

FlowerClassificationObject detection

Pollination is critical for crop development, especially those essential for subsistence. This study addresses the pollination challenges faced by Actinidia, a dioecious plant characterized by female and male flowers on separate plants. Despite the high protein content of pollen, the absence of nectar in kiwifruit flowers poses difficulties in attracting pollinators. Consequently, there is a growing interest in using artificial intelligence and robotic solutions to enable pollination even in unfavourable conditions. These robotic solutions must be able to accurately detect flowers and discern their genders for precise pollination operations. Specifically, upon identifying female Actinidia flowers, the robotic system should approach the stigma to release pollen, while male Actinidia flowers should target the anthers to collect pollen. We identified two primary research gaps: (1) the lack of gender-based flower detection methods and (2) the underutilisation of contemporary deep learning models in this domain. To address these gaps, we evaluated the performance of four pretrained models (YOLOv8, YOLOv5, RT-DETR and DETR) in detecting and determining the gender of Actinidia flowers. We outlined a comprehensive methodology and developed a dataset of manually annotated flowers categorized into two classes based on gender. Our evaluation utilised k-fold cross-validation to rigorously test model performance across diverse subsets of the dataset, addressing the limitations of conventional data splitting methods. DETR provided the most balanced overall performance, achieving precision, recall, F1 score and mAP of 89%, 97%, 93% and 94%, respectively, highlighting its robustness in managing complex detection tasks under varying conditions. These findings underscore the potential of deep learning models for effective gender-specific detection of Actinidia flowers, paving the way for advanced robotic pollination systems.

Why it matches plant phenotyping methodsキウイフルーツ花の画像検出と雌雄判定という植物器官の状態推定手法を開発・比較検証し、注釈付きデータセットと交差検証による性能評価も行っているため、フェノタイピング手法が中心である。

abstractwe evaluated the performance of four pretrained models (YOLOv8, YOLOv5, RT-DETR and DETR) in detecting and determining the gender of Actinidia flowers.
Reproduction assets foundThe paper's authors publicly deposited their gender-annotated Actinidia flower image dataset (augmented version) on Zenodo, explicitly linked in the Data availability statement. No author analysis code or trained model checkpoints are stated as publicly available.
Dataset · publicThe data presented in this study are openly available in the digital repository Zenodo: Actinidia chinensis cv. ’Hayward’ Flower Dataset 2024 (augmented version)— https://doi.org/10.5281/zenodo.13692222Open asset ↗Zenodo · 10.5281/zenodo.13692222lines:226-255
Code / dataset availability confirmedOpenAlex · checked 13 Sept 2026
Published16 Oct 2024Remote SensingCited by 6 · OpenAlex ↗

Influence of Structure from Motion Algorithm Parameters on Metrics for Individual Tree Detection Accuracy and Precision

Aerial / UAVField / plotPhotogrammetry / SfM / MVSStem / branchWhole plant / canopy / plot / fieldObject detectionArchitecture / morphology / geometryPlant / canopy height

Uncrewed aerial system (UAS) structure from motion (SfM) monitoring strategies for individual trees has rapidly expanded in the early 21st century. It has become common for studies to report accuracies for individual tree heights and DBH, along with stand density metrics. This study evaluates individual tree detection and stand basal area accuracy and precision in five ponderosa pine sites against the range of SfM parameters in the Agisoft Metashape, Pix4DMapper, and OpenDroneMap algorithms. The study is designed to frame UAS-SfM individual tree monitoring accuracy in the context of data processing and storage demands as a function of SfM algorithm parameter levels. Results show that when SfM algorithms are properly tuned, differences between software types are negligible, with Metashape providing a median F-score improvement over OpenDroneMap of 0.02 and PIX4DMapper of 0.06. However, tree extraction performance varied greatly across algorithm parameters, with the greatest extraction rates typically coming from parameters causing increased density in dense point clouds and minimal point cloud filtering. Transferring UAS-SfM forest monitoring into management will require tradeoffs between accuracy and efficiency. Our analysis shows that a one-step reduction in dense point cloud quality saves 77–86% in point cloud processing time without decreasing tree extraction (F-score) or basal area precision using Metashape and PIX4DMapper but the same parameter change for OpenDroneMap caused a ~5% loss in precision. Providing reproducible processing strategies is a vital step in successfully transferring these technologies into usage as management tools.

Why it matches plant phenotyping methodsUAS-SfMの処理パラメータと複数ソフトウェアを比較し、個体樹の抽出精度、樹高・DBH、林分断面積を評価する技術検証が中心である。

abstractThis study evaluates individual tree detection and stand basal area accuracy and precision in five ponderosa pine sites against the range of SfM parameters in the Agisoft Metashape, Pix4DMapper, and OpenDroneMap algorithms.
Reproduction assets foundThe paper's Data Availability Statement explicitly publishes the project's data and analysis source code to a public GitHub repository, which qualifies as a paper-specific public code asset. No separate phenotype dataset deposit is stated beyond this repository.
Code · publicData Availability Statement: Data and analysis source code for this project has been published to the public domain at: https://github.com/georgewoolsey/uas_sfm_tree_detection.Open asset ↗georgewoolsey/uas_sfm_tree_detectionpdf-page:19 lines:1-56
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published16 Oct 2024Cited by 1 · OpenAlex ↗

Identification of Plant Diseases in Jordan Using Convolutional Neural Networks

Brassica vegetablesCucumberLettuceTomatoField / plotWhole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

Abstract In the realm of global food security, plants serve as the primary source of sustenance. However, plant diseases pose a significant threat to this security. The process of diagnosing these diseases forms the bedrock of disease control efforts. The precision and expediency of these diagnoses wield substantial influence over disease management and the consequent reduction of economic losses. Conversely, incorrect diagnoses can render interventions ineffective, leading to agricultural crop deterioration and compounding economic hardships for both farmers and their respective nations. This research endeavors to diagnose the prevalent crops in Jordan, as identified by the Jordanian Department of Statistics for the year 2019. These crops encompass four key agricultural varieties: cucumbers, tomatoes, lettuce, and cabbage. To facilitate this, a novel dataset known as "Jordan 22" was meticulously curated. Jordan 22 was painstakingly compiled through the collection of images featuring both diseased and healthy plants, captured within the confines of Jordanian farms. These images underwent meticulous classification by a panel of three agricultural specialists, well-versed in plant disease identification and prevention. The Jordan 22 dataset comprises a substantial size, amounting to 3210 images. Following the compilation of this dataset, a series of preprocessing steps were executed. These encompassed the standardization of image backgrounds and the uniformization of image dimensions. Furthermore, image augmentation techniques were applied to the dataset to expand its diversity. Subsequently, a deep learning model, the Convolutional Neural Network (CNN), was meticulously trained on the augmented dataset. The results yielded by the CNN were nothing short of remarkable, with a test accuracy rate reaching an impressive 0.9712. Optimal performance was observed when images were resized to 256x256 dimensions, and max pooling was employed in lieu of average pooling within the pooling layer. Furthermore, the initial convolutional layer was set at a size of 32, with subsequent convolutional layers standardized at 128 in size. In conclusion, this research represents a pivotal step towards enhancing plant disease diagnosis and, by extension, global food security. Through the creation of the Jordan 22 dataset and the meticulous training of a CNN model, we have achieved substantial accuracy in disease detection, paving the way for more effective disease management strategies in agriculture.

Why it matches plant phenotyping methods植物画像から病害状態を推定するCNNとデータセットを開発・評価しており、植物フェノタイピング手法が研究の中心である。

abstractThis research endeavors to diagnose the prevalent crops in Jordan
Reproduction assets foundThe paper's Jordan22 plant disease image dataset (2310 RGB leaf images of cucumber, tomato, cabbage, and lettuce collected in Jordan and expert-classified) is explicitly stated as openly available on the authors' public GitHub repository. No separate analysis code or trained model checkpoint is explicitly deposited.
Dataset · publicThe data that support the findings of this study are openly available in [Jordan22_Dataset] at [https://github.com/shahd1995913/Jordan22_Dataset], reference number [17].Open asset ↗Jordan22_Datasetpdf-page:25 lines:1-40
Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Published9 Oct 2024arXivCited by 0 · OpenAlex ↗

NeRF-Accelerated Ecological Monitoring in Mixed-Evergreen Redwood Forest

NeRF / 3D Gaussian SplattingLiDAR / point cloudStem / branchMorphology / geometry measurement2D/3D reconstructionArchitecture / morphology / geometryBiomass / plant weight

Forest mapping provides critical observational data needed to understand the dynamics of forest environments. Notably, tree diameter at breast height (DBH) is a metric used to estimate forest biomass and carbon dioxide sequestration. Manual methods of forest mapping are labor intensive and time consuming, a bottleneck for large-scale mapping efforts. Automated mapping relies on acquiring dense forest reconstructions, typically in the form of point clouds. Terrestrial laser scanning (TLS) and mobile laser scanning (MLS) generate point clouds using expensive LiDAR sensing, and have been used successfully to estimate tree diameter. Neural radiance fields (NeRFs) are an emergent technology enabling photorealistic, vision-based reconstruction by training a neural network on a sparse set of input views. In this paper, we present a comparison of MLS and NeRF forest reconstructions for the purpose of trunk diameter estimation in a mixed-evergreen Redwood forest. In addition, we propose an improved DBH-estimation method using convex-hull modeling. Using this approach, we achieved 1.68 cm RMSE, which consistently outperformed standard cylinder modeling approaches. Our code contributions and forest datasets are freely available at https://github.com/harelab-ucsc/RedwoodNeRF.

Why it matches plant phenotyping methodsNeRFおよびMLSによる森林再構成から樹木DBHを推定し、凸包モデルによる推定法を提案・比較検証しているため、植物形質取得手法が中心です。

abstractIn this paper, we present a comparison of MLS and NeRF forest reconstructions for the purpose of trunk diameter estimation in a mixed-evergreen Redwood forest.
Reproduction assets foundThe authors explicitly state their code contributions and forest datasets (SLAM and NeRF reconstructions used for DBH estimation) are freely available in a public GitHub repository. The other URLs are a cited third-party tool (NeRFCapture) and a background reference (USDA aerial survey), neither of which is a paper-own
Dataset · publicOur code contributions and forest datasets are freely available at https://github.com/harelab-ucsc/RedwoodNeRF .Open asset ↗harelab-ucsc/RedwoodNeRFlines:1-52
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published9 Oct 2024Plant phenomics (Washington, D.C.)Cited by 19 · OpenAlex ↗

GSP-AI: An AI-Powered Platform for Identifying Key Growth Stages and the Vegetative-to-Reproductive Transition in Wheat Using Trilateral Drone Imagery and Meteorological Data.

WheatField / plotWhole plant / canopy / plot / fieldClassificationGrowth / development / phenology

Wheat ( Triticum aestivum ) is one of the most important staple crops worldwide. To ensure its global supply, the timing and duration of its growth cycle needs to be closely monitored in the field so that necessary crop management activities can be arranged in a timely manner. Also, breeders and plant researchers need to evaluate growth stages (GSs) for tens of thousands of genotypes at the plot level, at different sites and across multiple seasons. These indicate the importance of providing a reliable and scalable toolkit to address the challenge so that the plot-level assessment of GS can be successfully conducted for different objectives in plant research. Here, we present a multimodal deep learning model called GSP-AI, capable of identifying key GSs and predicting the vegetative-to-reproductive transition (i.e., flowering days) in wheat based on drone-collected canopy images and multiseasonal climatic datasets. In the study, we first established an open Wheat Growth Stage Prediction (WGSP) dataset, consisting of 70,410 annotated images collected from 54 varieties cultivated in China, 109 in the United Kingdom, and 100 in the United States together with key climatic factors. Then, we built an effective learning architecture based on Res2Net and long short-term memory (LSTM) to learn canopy-level vision features and patterns of climatic changes between 2018 and 2021 growing seasons. Utilizing the model, we achieved an overall accuracy of 91.2% in identifying key GS and an average root mean square error (RMSE) of 5.6 d for forecasting the flowering days compared with manual scoring. We further tested and improved the GSP-AI model with high-resolution smartphone images collected in the 2021/2022 season in China, through which the accuracy of the model was enhanced to 93.4% for GS and RMSE reduced to 4.7 d for the flowering prediction. As a result, we believe that our work demonstrates a valuable advance to inform breeders and growers regarding the timing and duration of key plant growth and development phases at the plot level, facilitating them to conduct more effective crop selection and make agronomic decisions under complicated field conditions for wheat improvement.

Why it matches plant phenotyping methodsドローン画像と気象データからコムギの生育ステージおよび開花日を推定するモデルを開発・検証し、公開データセットも構築しており、表現型取得・推定手法が研究の中心である。

abstractwe present a multimodal deep learning model called GSP-AI, capable of identifying key GSs and predicting the vegetative-to-reproductive transition (i.e., flowering days) in wheat based on drone-collected canopy images and multiseasonal climatic datasets.
Reproduction assets foundThe paper's Data Availability statement explicitly deposits the authors' GSP-AI source code and the open WGSP dataset (70,410 annotated drone/smartphone canopy images with climatic data) on GitHub, plus the AirMeasurer plot-segmentation platform code. Both are paper-specific, public, and actionable.
Code · publicSource code, WGSP, and other datasets supporting the results presented in this article are available at https://Github.com/The-Zhou-Lab/GSP-AI/releases .Open asset ↗The-Zhou-Lab/GSP-AIlines:308-308
Code / dataset availability confirmedOpenAlex · Europe PMC · bioRxiv · checked 15 Sept 2026
Published5 Oct 2024bioRxiv (Cold Spring Harbor Laboratory)Cited by 3 · OpenAlex ↗

The FIP 1.0 Data Set: Highly Resolved Annotated Image Time Series of 4,000 Wheat Plots Grown in Six Years

WheatField / plotSeed / grainWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenologyPigment / colour / senescencePlant / canopy heightYield / yield components

Abstract Background Understanding genotype-environment interactions of plants is crucial for crop improvement, yet limited by the scarcity of quality phenotyping data. This data note presents the Field Phenotyping Platform 1.0 data set, a comprehensive resource for winter wheat research that combines imaging, trait, environmental, and genetic data. Findings We provide time series data for more than 4,000 wheat plots, including aligned high-resolution image sequences totaling more than 153,000 aligned images across six years. Measurement data for eight key wheat traits is included, namely canopy cover values, plant heights, wheat head counts, senescence ratings, heading date, final plant height, grain yield, and protein content. Genetic marker information and environmental data complement the time series. Data quality is demonstrated through heritability analyses and genomic prediction models, achieving accuracies aligned with previous research. Conclusions This extensive data set offers opportunities for advancing crop modeling and phenotyping techniques, enabling researchers to develop novel approaches for understanding genotype-environment interactions, analyzing growth dynamics, and predicting crop performance. By making this resource publicly available, we aim to accelerate research in climate-adaptive agriculture and foster collaboration between plant science and machine learning communities.

Why it matches plant phenotyping methods高解像度画像時系列と複数の植物形質を含む大規模な公開圃場フェノタイピングデータセットであり、再利用可能なフェノタイピング基盤・ベンチマークとして中心的です。

abstractThis data note presents the Field Phenotyping Platform 1.0 data set, a comprehensive resource for winter wheat research that combines imaging, trait, environmental, and genetic data.
Reproduction assets foundThis data note directly publishes its own phenotyping measurements and image time series: the FIP 1.0 dataset (images, aligned image sequences, eight wheat traits, environmental and marker data) is publicly available on the ETH Research Collection and Hugging Face, and the authors' analysis/processing code is publicly,
Dataset · publicn License: GNU GPL v3 Data Set Compilation Project name: fip1-dataset Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-dataset Operating system(s): Platform independent Programming language: Python License: GNU GPL v3 Data Availability • Data Repository: http://doi.org/20.500.11850/697773 • Hugging Face Data set: https://huggingface.co/datasets/mikeboss/FIP1 • Public GABI marker data repository (also integrated in main Data Repository and Hugging Face Data set): https://doi.org/10.5061/dryad.n02v6wwzc • Private Agroscope marker data repository: Confidential (Con- tact: Boulos Chalhoub, boulos.chalhoub@agroscope.admin.ch). This repository contains marker data (Illumina InfiniumOpen asset ↗mikeboss/FIP1pdf-raw-page:7 lines:1-110
Dataset · publicn with FAIR principles [26]: • Findable: This publication and the Hugging Face data set card (https://doi.org/10.57967/hf/3191) provide detailed meta- data and a comprehensive description of the data set’s contents, making it discoverable to researchers. • Accessible: The data is hosted on the Research Collection of ETH Zurich (https://doi.org/20.500.11850/697773), a reliable and openly accessible data storage. • Interoperable: The use of the open-source Hugging Face datasets [27] package makes it easy to use and export to differ- ent formats. The data is fully MIAPPE v1.1 [28] conform. Given the shared genotypes the data set can be used to enhance the data by Gogna et al. [13] by 6 envOpen asset ↗20.500.11850/697773pdf-raw-page:2 lines:84-130
Code · publiclly aggregating the derived data into the final data set using the fip1-dataset repository. In addition, the data set can be recreated using the fip1-dataset repository from the derived data that is freely available in the ETH research collection. Trait Data Compilation Project name: FIP 1.0 Data Set - Traits Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip-1.0-data-set-traits Operating system(s): Platform independent Programming language: R, Python License: GNU GPL v3 Image Data Alignment Project name: fip1-alignment Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-alignment Operating system(s): Platform independent Programming language: Python License: GNU GPL Open asset ↗fip-1.0-data-set-traitspdf-raw-page:7 lines:1-110
Code · publicresearch collection. Trait Data Compilation Project name: FIP 1.0 Data Set - Traits Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip-1.0-data-set-traits Operating system(s): Platform independent Programming language: R, Python License: GNU GPL v3 Image Data Alignment Project name: fip1-alignment Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-alignment Operating system(s): Platform independent Programming language: Python License: GNU GPL v3 Data Set Compilation Project name: fip1-dataset Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-dataset Operating system(s): Platform independent Programming language: Python License: GNU GPL v3 Data AvailabiOpen asset ↗fip1-alignmentpdf-raw-page:7 lines:1-110
Code · publicramming language: R, Python License: GNU GPL v3 Image Data Alignment Project name: fip1-alignment Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-alignment Operating system(s): Platform independent Programming language: Python License: GNU GPL v3 Data Set Compilation Project name: fip1-dataset Project home page: https://gitlab.ethz.ch/crop_phenotyping/fip1-dataset Operating system(s): Platform independent Programming language: Python License: GNU GPL v3 Data Availability • Data Repository: http://doi.org/20.500.11850/697773 • Hugging Face Data set: https://huggingface.co/datasets/mikeboss/FIP1 • Public GABI marker data repository (also integrated in main Data Repository and HOpen asset ↗fip1-datasetpdf-raw-page:7 lines:1-110
Code / dataset availability confirmedCrossref · Europe PMC · checked 13 Sept 2026
Published1 Oct 2024Data in BriefCited by 14 · OpenAlex ↗

Comprehensive smart smartphone image dataset for plant leaf disease detection and freshness assessment from Bangladesh vegetable fields

Brassica vegetablesCucumberEggplant / aubergineTomatoField / plotLeafObject detectionStress / disease detectionDisease symptoms / severityYield / yield components

Bangladesh's agricultural landscape is significantly influenced by vegetable cultivation, which substantially enhances nutrition, the economy, and food security in the nation. Millions of people rely on vegetable production for their daily sustenance, generating considerable income for numerous farmers. However, leaf diseases frequently compromise the yield and quality of vegetable crops. Plant diseases are a common impediment to global agricultural productivity, adversely affecting crop quality and yield, leading to substantial economic losses for farmers. Early detection of plant leaf diseases is crucial for improving cultivation and vegetable production. Common diseases such as Bacterial Spot, Mosaic Virus, and Downy Mildew often reduce vegetable plant cultivation and severely impact vegetable production and the food economy. Consequently, many farmers in Bangladesh struggle to identify the specific diseases, incurring significant losses. This dataset contains 12,643 images of widely grown crops in Bangladesh, facilitating the identification of unhealthy leaves compared to healthy ones. The dataset includes images of vegetable leaves such as Bitter Gourd (2223 images), Bottle Gourd (1803 images), Eggplants (2944 images), Cauliflowers (1598 images), Cucumbers (1626 images), and Tomatoes (2449 images). Each vegetable class encompasses several common diseases that affect cultivation. By identifying early leaf diseases, this dataset will be invaluable for farmers and agricultural researchers alike.

Why it matches plant phenotyping methods植物葉の画像から健康状態と病徴を識別するデータセットを提供しており、植物の病害状態を観測する再利用可能な画像ベースのフェノタイピング資源が中心です。

abstractThis dataset contains 12,643 images of widely grown crops in Bangladesh, facilitating the identification of unhealthy leaves compared to healthy ones.
Reproduction assets foundThe paper is a Data in Brief article describing a public smartphone image dataset of vegetable leaf diseases hosted on Mendeley Data, with an explicit direct URL matching an allowed URL.
Dataset · publicRepository name: Mendeley Data Data identification number: DOI: 10.17632/n67gctmjyj.3 Direct URL to data: https://data.mendeley.com/datasets/n67gctmjyj/3Open asset ↗Mendeley Data · 10.17632/n67gctmjyj.3lines:1-48
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 7 Sept 2026
Published18 Sept 2024Frontiers in Artificial IntelligenceCited by 3 · OpenAlex ↗

Investigating the contribution of image time series observations to cauliflower harvest-readiness prediction.

Brassica vegetablesWhole plant / canopy / plot / fieldClassificationGrowth / time-series analysisGrowth / development / phenology

Cauliflower cultivation is subject to high-quality control criteria during sales, which underlines the importance of accurate harvest timing. Using time series data for plant phenotyping can provide insights into the dynamic development of cauliflower and allow more accurate predictions of when the crop is ready for harvest than single-time observations. However, data acquisition on a daily or weekly basis is resource-intensive, making selection of acquisition days highly important. We investigate which data acquisition days and development stages positively affect the model accuracy to get insights into prediction-relevant observation days and aid future data acquisition planning. We analyze harvest-readiness using the cauliflower image time series of the GrowliFlower dataset. We use an adjusted ResNet18 classification model, including positional encoding of the data acquisition dates to add implicit information about development. The explainable machine learning approach GroupSHAP analyzes time points' contributions. Time points with the lowest mean absolute contribution are excluded from the time series to determine their effect on model accuracy. Using image time series rather than single time points, we achieve an increase in accuracy of 4%. GroupSHAP allows the selection of time points that positively affect the model accuracy. By using seven selected time points instead of all 11 ones, the accuracy improves by an additional 4%, resulting in an overall accuracy of 89.3%. The selection of time points may therefore lead to a reduction in data collection in the future.

Why it matches plant phenotyping methodsカリフラワー画像時系列を用いた収穫適期という植物状態の推定手法を開発・評価し、データ取得時点の選択とモデル精度を検証しているため、フェノタイピング手法が中心である。

abstractUsing time series data for plant phenotyping can provide insights into the dynamic development of cauliflower and allow more accurate predictions of when the crop is ready for harvest than single-time observations.
Reproduction assets foundThe paper analyzes the GrowliFlower cauliflower UAV image time series dataset and provides a public data availability link to the dataset metadata on phenoroam.phenorob.de. No author analysis code or trained model checkpoint is explicitly deposited.
Dataset · publicPublicly available datasets were analyzed in this study. This data can be found at: https://phenoroam.phenorob.de/geonetwork/srv/eng/catalog.search#/metadata/cb328232-31f5-4b84-a929-8e1ee551d66a .Open asset ↗phenoroam.phenorob.de · cb328232-31f5-4b84-a929-8e1ee551d66alines:386-397
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published18 Sept 2024Data in briefCited by 7 · OpenAlex ↗

Lidar-derived structural-complexity data across four experimental forests.

Aerial / UAVField / plotLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detectionCalibration / preprocessingSegmentationArchitecture / morphology / geometryPlant / canopy height

Structural complexity refers to the three-dimensional arrangement and variability of both biotic and abiotic components of an ecosystem. Metrics that characterize structural complexity are often used to manage various aspects of ecosystem function, such as light transmittance, wildlife habitat, and biological diversity. Additionally, these metrics aid in evaluating resilience to disturbance events, including hurricanes, bark-beetle outbreaks, and wildfire. Recent advances in wildland fire modelling have facilitated the integration of forest structural complexity metrics into the QUIC-Fire model, enabling real-time prediction of fire spread and behaviour by simulating interactions between fire, weather, topography, and forest structure. While QUIC-Fire is designed to be highly adaptable, model performance depends on the availability and accuracy of local data inputs. Expanding the model's usability across different regions can be facilitated by the availability of more comprehensive and high-quality data. Thus, the primary goal behind the data products we developed was to establish a basis for collaborative research across various disciplines, particularly within the focal areas of the Southern Research Station, such as forestry, wildland fire, hydrology, soil science, and cultural resources at Bent Creek, Coweeta, Escambia, and Hitchiti Experimental Forests (EFs). Airborne laser scanning (ALS) was used to collect point-cloud data for each EF during the leaf-off season to minimize interference from foliage. Subsequent processing of the raw lidar data involved outlier detection and filtering, ground and non-ground classification, and the computation of a variety of metrics representing various aspects of topography and forest structure at both the pixel-level and the tree-level. Pixel-level topographic data products include: digital elevation model (DEM), slope, aspect, topographic position index (TPI), topographic roughness index (TRI), roughness, and flow direction. Forest structural-complexity metrics include canopy height, foliar height diversity (FHD), vertical distribution ratio (VDR), canopy rugosity, crown relief ratio (CRR), understory complexity index (UCI), vertical complexity index (VCI), canopy cover, mean vegetation height, and the standard deviation of vegetation height. Tree-level data products were computed from the point cloud using multiple algorithms to perform individual tree detection (ITD) and individual tree segmentation (ITS). The datasets have been harmonized and are openly accessible through the USDA Forest Service Research Data Archive.

Why it matches plant phenotyping methods航空レーザースキャンから樹冠高、植生高、樹冠構造、個体樹木を抽出した再利用可能なデータセットであり、植物の構造形質取得と処理が中心です。

abstractAirborne laser scanning (ALS) was used to collect point-cloud data for each EF during the leaf-off season
Reproduction assets foundThis Data in Brief article describes its own openly archived dataset: lidar-derived forest structural-complexity metrics (raster and vector products, including tree detection/segmentation outputs) for four experimental forests, deposited in the USFS Research Data Archive (RDS-2024-0019) with R processing code in theSup
Dataset · public−83.450054 Coweeta Experimental Forests: 31.007539, −87.078571 Escambia Experimental Forests: 33.057078, −83.679620 Hitchiti Experimental Forest: 35.484250, −82.633346 Data accessibility Repository name: US Forest Service Research Data Archive Data identification number: https://doi.org/10.2737/RDS-2024-0019 Direct URL to data: https://www.fs.usda.gov/rds/archive/catalog/RDS-2024-0019 Raw ALS point-cloud data are located at https://app.box.com/s/4s3412g8mtky0hb6wb63a44epv08c22o Related research article none. 1. Value of the Data •Open asset ↗US Forest Service Research Data Archive · RDS-2024-0019lines:1-50
Dataset · publicarch Station. Bent Creek Experimental Forests: 35.050580, −83.450054 Coweeta Experimental Forests: 31.007539, −87.078571 Escambia Experimental Forests: 33.057078, −83.679620 Hitchiti Experimental Forest: 35.484250, −82.633346 Data accessibility Repository name: US Forest Service Research Data Archive Data identification number: https://doi.org/10.2737/RDS-2024-0019 Direct URL to data: https://www.fs.usda.gov/rds/archive/catalog/RDS-2024-0019 Raw ALS point-cloud data are located at https://app.box.com/s/4s3412g8mtky0hb6wb63a44epv08c22o Related research article none. 1. Value of the Data •Open asset ↗US Forest Service Research Data Archive · RDS-2024-0019lines:1-50
Dataset · publicean crown diameter. Additionally, the generalized additive model (GAM) that was developed from the inventory data was used to predict bole height at the tree-level. Limitations The size of the raw point-cloud data precluded storage on the USFS Research Data Archive. Therefore, this data is accessible for download from box.com ( https://app.box.com/s/4s3412g8mtky0hb6wb63a44epv08c22o ). Additionally, the volume of the point-cloud data may pose computational limitations. Ethics Statement The authors have read and follow the ethical requirements for publication in Data in Brief and confirm that the current work does not involve human subjects, animal experiments, or any data collected from sociaOpen asset ↗lines:420-434
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 Sept 2024Data in briefCited by 4 · OpenAlex ↗

High-resolution image dataset for the automatic classification of phenological stage and identification of racemes in Urochloa spp. hybrids.

RGB / grayscalePanicle / ear / spikeWhole plant / canopy / plot / fieldClassificationObject detectionGrowth / development / phenology

Urochloa grasses are widely used forages in the Neotropics and are gaining importance in other regions due to their role in meeting the increasing global demand for sustainable agricultural practices. High-throughput phenotyping (HTP) is important for accelerating Urochloa breeding programs focused on improving forage and seed yield. While RGB imaging has been used for HTP of vegetative traits, the assessment of phenological stages and seed yield using image analysis remains unexplored in this genus. This work presents a dataset of 2,400 high-resolution RGB images of 200 Urochloa hybrid genotypes, captured over seven months and covering both vegetative and reproductive stages. Images were manually labelled as vegetative or reproductive, and a subset of 255 reproductive stage images were annotated to identify 22,340 individual racemes. This dataset enables the development of machine learning and deep learning models for automated phenological stage classification and raceme identification, facilitating HTP and accelerated breeding of Urochloa spp. hybrids with high seed yield potential.

Why it matches plant phenotyping methods植物のフェノロジー段階と穂状花序を画像から識別するための高解像度データセットであり、植物表現型取得・抽出を中心とする研究。

abstractThis work presents a dataset of 2,400 high-resolution RGB images of 200 Urochloa hybrid genotypes, captured over seven months and covering both vegetative and reproductive stages.
Reproduction assets foundThe paper is itself a data descriptor whose core asset is a public Harvard Dataverse dataset of 2,400 RGB images of Urochloa hybrids with phenological stage labels and COCO-format raceme polygon annotations, directly reproducing the paper's phenotyping data.
Dataset · publict diffuser to ensure uniform lighting. Data source location Institution: Alliance Bioversity International & CIAT. City: Palmira, Valle del Cauca. Country: Colombia. Geolocalization: 3°29′N, 76°21′W . Data accessibility Repository name: Harvard Dataverse Data identification number: doi.org/10.7910/dvn/u0kl6y Direct URL to data: https://doi.org/10.7910/dvn/u0kl6y Instructions for accessing these data: The dataset [ 1 ] is licensed under the Creative Commons Attribution 4.0 International, which allows using, sharing, adapting, distribution and reproduction in any medium or format if attribution is given to the creator. Related research article None . 1 Value of the Data • The dataset offOpen asset ↗Harvard Dataverselines:1-58
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published13 Sept 2024Tenth International Conference on Remote Sensing and Geoinformation of the Environment (RSCy2024)Cited by 1 · OpenAlex ↗

Estimation of crop yield using deep learning for precision agriculture

MangoWheatFruitPanicle / ear / spikeCountingObject detectionYield / yield components

Precision agriculture is the application of correct amount of fertilizers and water pesticide to achieve higher agricultural productivity. Furthermore, under the framework of precision agriculture is the automated estimation of yield with advanced technologies including Artificial Intelligence (AI) and Remote Sensing (RS). The use of RS has advanced crop yield estimations and predictions in recent years. However, to validate RS-based models it is important to perform in-situ exercises such as fruit counting, which is a time-consuming task that increases the production costs. Drones, robots, and in-situ cameras in combination with AI algorithms are widely used to efficiently address these issues. The recent advancement in computational resources and power available has enabled the utilization of Deep Learning AI models. One of the best-performing models for object detection is the You-Only-Look-Once (YOLO). In this study, the YOLOv5s is used for object detection, which is the second smallest and fastest YOLOv5 architecture, on two different benchmark datasets collected from AgML. The first dataset consists of 1730 images of mango trees in Australia during night, and the second dataset consists of 6512 images of wheat heads collected from different regions around the world. The main objective of this work is to demonstrate the capabilities of light AI models for object detection and to evaluate their performance, which will serve as a benchmark for future comparison with the on-board environment.

Why it matches plant phenotyping methods植物器官の検出・カウントによる収量推定を対象とし、YOLOv5sの性能評価とベンチマーク化が主目的であるため、計算画像フェノタイピング手法として採用。

abstractIn this study, the YOLOv5s is used for object detection, which is the second smallest and fastest YOLOv5 architecture, on two different benchmark datasets collected from AgML.
Reproduction assets foundThe paper evaluates YOLOv5s on two public benchmark datasets. The MangoYOLO dataset is explicitly cited with public access URLs and was directly used for the paper's mango yield-estimation experiments, qualifying as a paper-specific public asset. The Global Wheat Head Detection dataset is also used but its Zenodo URL (
Dataset · publicAnand Koirala, C McCarthy, Kerry Walsh, and Z Wang, ‘MangoYOLO data set’. Central Queensland University, 2021. Accessed: May 23, 2024. [Online]. Available: http://hdl.handle.net/10018/1261224, https://researchdata.edu.au/mangoyolo-setOpen asset ↗pdf-page:7 lines:1-50
Code / dataset availability confirmedCrossref · checked 13 Sept 2026
Published11 Sept 2024AgricultureCited by 10 · OpenAlex ↗

Cotton Disease Recognition Method in Natural Environment Based on Convolutional Neural Network

CottonField / plotClassificationDisease symptoms / severity

As an essential component of the global economic crop, cotton is highly susceptible to the impact of diseases on its yield and quality. In recent years, artificial intelligence technology has been widely used in cotton crop disease recognition, but in complex backgrounds, existing technologies have certain limitations in accuracy and efficiency. To overcome these challenges, this study proposes an innovative cotton disease recognition method called CANnet, and we independently collected and constructed an image dataset containing multiple cotton diseases. Firstly, we introduced the innovatively designed Reception Field Space Channel (RFSC) module to replace traditional convolution kernels. This module combines dynamic receptive field features with traditional convolutional features to effectively utilize spatial channel attention, helping CANnet capture local and global features of images more comprehensively, thereby enhancing the expressive power of features. At the same time, the module also solves the problem of parameter sharing. To further optimize feature extraction and reduce the impact of spatial channel attention redundancy in the RFSC module, we connected a self-designed Precise Coordinate Attention (PCA) module after the RFSC module to achieve redundancy reduction. In the design of the classifier, CANnet abandoned the commonly used MLP in traditional models and instead adopted improved Kolmogorov Arnold Networks-s (KANs) for classification operations. KANs technology helps CANnet to more finely utilize extracted features for classification tasks through learnable activation functions. This is the first application of the KAN concept in crop disease recognition and has achieved excellent results. To comprehensively evaluate the performance of CANnet, we conducted extensive experiments on our cotton disease dataset and a publicly available cotton disease dataset. Numerous experimental results have shown that CANnet outperforms other advanced methods in the accuracy of cotton disease identification. Specifically, on the self-built dataset, the accuracy reached 96.3%; On the public dataset, the accuracy reached 98.6%. These results fully demonstrate the excellent performance of CANnet in cotton disease identification tasks.

Why it matches plant phenotyping methods綿花の病徴を画像から認識するCNN手法を開発し、自作および公開データセットで性能検証しており、植物病害状態の画像ベース表現型取得が中心である。

abstractthis study proposes an innovative cotton disease recognition method called CANnet
Reproduction assets foundThe paper evaluates CANnet on a self-built Xinjiang cotton disease dataset (no public availability statement) and on a public Kaggle cotton disease dataset explicitly cited with a URL. No author code, models, or supplementary data deposits are mentioned.
Dataset · public29. Dhamodharan. Cotton Plant Disease. 2023. Available online: https://www.kaggle.com/datasets/dhamur/cotton-plant-disease (accessed on 6 May 2024).Open asset ↗Kaggle · dhamur/cotton-plant-diseasepdf-page:21 lines:1-56
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Sept 2024Data in briefCited by 20 · OpenAlex ↗

BDPapayaLeaf: A dataset of papaya leaf for disease detection, classification, and analysis.

LeafClassificationSegmentationStress / disease detectionDisease symptoms / severity

Papaya is a popular vegetable and fruit in both developing and developed countries. Nonetheless, Bangladesh's agricultural landscape is significantly influenced by papaya cultivation. However, disease is a common impediment to papaya productivity, adversely affecting papaya quality and yield and leading to substantial economic losses for farmers. Research suggests that computer-aided disease diagnosis and machine learning (ML) models can improve papaya production by detecting and classifying diseases. In this line, a dataset of papaya is required to diagnose the disease. Moreover, like many other fruits, papaya disease may vary from country to country. Therefore, the country-based papaya disease dataset is required. In this study, a papaya dataset is collected from Dhaka, Bangladesh. This dataset contains 2159 original images from five classes, including the healthy control class and four papaya leaf diseases: Anthracnose, Bacterial Spot, Curl, and Ring spot. Besides the original images, the dataset contains 210 annotated data for each of the five classes. The dataset contains two types of data: the whole image and the annotated image . The image will interest data scientists who apply disease detection through a convolutional neural network (CNN) and its variants. Furthermore, the annotated images, such as You Only Look Once (YOLO), U-Net, Mask R-CNN, and Single Shot Detection (SSD), will be helpful for semantic segmentation. Since firm-applicable AI devices and mobile and web applications are in demand, the dataset collected in this study will offer multiple options for integrating ML models into AI devices. In countries with weather and climate similar to Bangladesh, data scientists may use their dataset in that context.

Why it matches plant phenotyping methodsパパイヤ葉の病害状態を画像とアノテーションで記録したデータセット自体が中心であり、植物病害表現型の検出・分類・セグメンテーションに再利用可能な資源を提供している。

abstractThis dataset contains 2159 original images from five classes, including the healthy control class and four papaya leaf diseases: Anthracnose, Bacterial Spot, Curl, and Ring spot.
Reproduction assets foundThe paper is a Data in Brief article describing the BDPapayaLeaf dataset of 2159 papaya leaf images (5 classes) with XML/TXT annotations, publicly deposited on Mendeley Data with a direct URL and DOI provided in the article.
Dataset · publicd location, and the annotations were saved in XML and TXT forms for usage with various models. Data Source Location City: Changao, Ashulia, Dhaka Country: Bangladesh Coordinates: 23° 53′ 2″ N and 90° 19′ 28″ E Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/p997fvf526.1 Direct URL to data: https://data.mendeley.com/datasets/p997fvf526/2 1 Value of the Data •Open asset ↗Mendeley Data · 10.17632/p997fvf526.1lines:1-49
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Sept 2024Data in briefCited by 54 · OpenAlex ↗

A comprehensive cotton leaf disease dataset for enhanced detection and classification.

CottonField / plotLeafClassificationStress / disease detectionDisease symptoms / severity

The creation and use of a comprehensive cotton leaf disease dataset offer significant benefits in agricultural research, precision farming, and disease management. This dataset enables the development of accurate machine learning models for early disease detection, reducing manual inspections and facilitating timely interventions. It serves as a benchmark for testing algorithms and training deep learning models, aiding in automated monitoring and decision support tools in precision agriculture. This leads to targeted interventions, reduced chemical use, and improved crop management. Global collaboration is fostered, contributing to the development of disease-resistant cotton varieties and effective management strategies, ultimately reducing economic losses and promoting sustainable farming. Field surveys conducted from October 2023 to January 2024 ensured meticulous image capture under diverse conditions. The images are categorized into eight classes, representing specific disease manifestations, pests, or environmental stress in cotton plants. The dataset comprises 2137 original images and 7000 augmented images, enhancing deep learning model training. The Inception V3 model demonstrated high performance, with an overall accuracy of 96.03 %. This underscores the dataset's potential in advancing automated disease detection in cotton agriculture.

Why it matches plant phenotyping methods綿花葉の病徴を画像で分類するデータセットを構築し、深層学習モデルのベンチマークとして評価しており、植物病害表現型の取得・解析が中心である。

abstractThe creation and use of a comprehensive cotton leaf disease dataset offer significant benefits in agricultural research, precision farming, and disease management.
Reproduction assets foundThe paper is a Data in Brief article describing the authors' own cotton leaf disease image dataset (SAR-CLD-2024), publicly deposited on Mendeley Data with a direct URL and DOI provided in the article.
Dataset · publicining and evaluating machine learning models aimed at accurately classifying and diagnosing cotton leaf diseases. Data source location The National Cotton Research Institute field in Gazipur, Dhaka, Bangladesh Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/b3jy2p6k8w.2 Direct URL to data: https://data.mendeley.com/datasets/b3jy2p6k8w/2 Related research article None 1. Value of the Data • The presence of diseases such as Cotton Leaf Curl Disease and leaf hopper in cotton plants poses significant challenges to farmers worldwide, leading to substantial yield losses, reduced crop quality, and economic hardships. Timely detection and effective management ofOpen asset ↗Mendeley Data · 10.17632/b3jy2p6k8w.2lines:1-44
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Sept 2024Data in briefCited by 2 · OpenAlex ↗

Comprehensive stomata image dataset of Sundarbans Mangrove and Ratargul Swamp forest tree species in Bangladesh.

Field / plotStomata / guard-cell complexClassificationObject detectionStomatal traits

Plants' leaf stomata are crucial for various scientific research, including identifying species, studying ecology, conserving ecosystems, improving agriculture, and advancing the field of deep learning. This dataset, containing 1083 images, encompasses 11 species from two distinct locations in Bangladesh: nine from the Sundarbans mangrove forest and two from the Ratargul Swamp Forest. It is a valuable tool for refining machine learning algorithms that specialize in detecting stomata and categorizing species accurately. Researchers can explore a deeper understanding of plant physiology, adaptation mechanisms, and environmental interactions by employing pattern recognition, deep learning, and feature extraction techniques. Additionally, this dataset could be a potential tool for enhancing research in macroscopic metamaterials, extending its impact beyond traditional biological studies into interdisciplinary fields of technology and material science.

Why it matches plant phenotyping methods気孔画像データセット自体が研究の中心で、画像から気孔を検出する再利用可能な表現型取得基盤を提供しているため。

abstractThis dataset, containing 1083 images, encompasses 11 species from two distinct locations in Bangladesh
Reproduction assets foundThis Data in Brief article deposits its own stomata image dataset (1083 images, 11 species) and stomatal trait measurements in Mendeley Data, with a direct public URL and DOI given in the Specifications Table.
Dataset · publicSiedentopf Trinocular Compound Microscope (AmScope T340A) was used to image stomata with AmScope camera software. Data source location Country: BangladeshForest: Sundarbans Mangrove Forest, Ratargul Swamp Forest Data accessibility Repository name: Mendeley DataData identification number: 10.17632/4brcwhmvyk.4Direct URL to data: https://data.mendeley.com/datasets/4brcwhmvyk/4 Related research article Dey, B., Ahmed, R., Ferdous, J., Haque, M.M.U., Khatun, R., Hasan, F.E., Uddin, S.N., 2023. Automated plant species identification from the stomata images using deep neural network: A study of selected mangrove and freshwater swamp forest tree species of Bangladesh. Ecol. Inform. 75, 102128.httpsOpen asset ↗Mendeley Data · 10.17632/4brcwhmvyk.4html-lines:1-37
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 7 Sept 2026
Published28 Aug 2024Plant PhenomicsCited by 13 · OpenAlex ↗

Deep Learning Methods Using Imagery from a Smartphone for Recognizing Sorghum Panicles and Counting Grains at a Plant Level

SorghumField / plotPanicle / ear / spikeSeed / grainWhole plant / canopy / plot / fieldCountingObject detectionSegmentationYield / biomass estimationFruit / seed / panicle traits

High-throughput phenotyping is the bottleneck for advancing field trait characterization and yield improvement in major field crops. Specifically for sorghum ( Sorghum bicolor L.), rapid plant-level yield estimation is highly dependent on characterizing the number of grains within a panicle. In this context, the integration of computer vision and artificial intelligence algorithms with traditional field phenotyping can be a critical solution to reduce labor costs and time. Therefore, this study aims to improve sorghum panicle detection and grain number estimation from smartphone-capture images under field conditions. A preharvest benchmark dataset was collected at field scale (2023 season, Kansas, USA), with 648 images of sorghum panicles retrieved via smartphone device, and grain number counted. Each sorghum panicle image was manually labeled, and the images were augmented. Two models were trained using the Detectron2 and Yolov8 frameworks for detection and segmentation, with an average precision of 75% and 89%, respectively. For the grain number, 3 models were trained: MCNN (multiscale convolutional neural network), TCNN-Seed (two-column CNN-Seed), and Sorghum-Net (developed in this study). The Sorghum-Net model showed a mean absolute percentage error of 17%, surpassing the other models. Lastly, a simple equation was presented to relate the count from the model (using images from only one side of the panicle) to the field-derived observed number of grains per sorghum panicle. The resulting framework obtained an estimation of grain number with a 17% error. The proposed framework lays the foundation for the development of a more robust application to estimate sorghum yield using images from a smartphone at the plant level.

Why it matches plant phenotyping methodsスマートフォン画像からソルガム穂の検出・分割と粒数推定を開発・検証しており、植物形質取得手法が研究の中心である。

abstractthis study aims to improve sorghum panicle detection and grain number estimation from smartphone-capture images under field conditions.
Reproduction assets foundThe paper's authors explicitly state that the code used to train, test, and analyze the data is publicly available on GitHub. The phenotype image datasets are only available upon request.
Code · publicThe code used to train, test, and analyze the data is available at https://github.com/GustavoSantiago113/Sorghum_Grain_Counter .Open asset ↗GustavoSantiago113/Sorghum_Grain_Counterlines:169-296
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published24 Aug 2024BiologyCited by 13 · OpenAlex ↗

i PhyDSDB: Phytoplasma Disease and Symptom Database.

Annotation / quality controlVisualization / data managementDisease symptoms / severity

Phytoplasmas are small, intracellular bacteria that infect a vast range of plant species, causing significant economic losses and impacting agriculture and farmers' livelihoods. Early and rapid diagnosis of phytoplasma infections is crucial for preventing the spread of these diseases, particularly through early symptom recognition in the field by farmers and growers. A symptom database for phytoplasma infections can assist in recognizing the symptoms and enhance early detection and management. In this study, nearly 35,000 phytoplasma sequence entries were retrieved from the NCBI nucleotide database using the keyword "phytoplasma" and information on phytoplasma disease-associated plant hosts and symptoms was gathered. A total of 945 plant species were identified to be associated with phytoplasma infections. Subsequently, links to symptomatic images of these known susceptible plant species were manually curated, and the Phytoplasma Disease Symptom Database ( i PhyDSDB) was established and implemented on a web-based interface using the MySQL Server and PHP programming language. One of the key features of i PhyDSDB is the curated collection of links to symptomatic images representing various phytoplasma-infected plant species, allowing users to easily access the original source of the collected images and detailed disease information. Furthermore, images and descriptive definitions of typical symptoms induced by phytoplasmas were included in i PhyDSDB. The newly developed database and web interface, equipped with advanced search functionality, will help farmers, growers, researchers, and educators to efficiently query the database based on specific categories such as plant host and symptom type. This resource will aid the users in comparing, identifying, and diagnosing phytoplasma-related diseases, enhancing the understanding and management of these infections.

Why it matches plant phenotyping methods植物の病徴画像と症状定義を体系的に収録し、植物病害状態の認識・診断に利用するデータベースとウェブインターフェースを開発した研究であり、病徴という植物状態の取得・参照基盤が中心です。

abstractSubsequently, links to symptomatic images of these known susceptible plant species were manually curated, and the Phytoplasma Disease Symptom Database ( i PhyDSDB) was established and implemented on a web-based interface using the MySQL Server and PHP programming language.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicwe established a database that consists of various phytoplasma diseases and their associated symptoms, and we implemented it on a web-based interface ( https://plantpathology.ba.ars.usda.gov/iphydsdb/iphydsdb.html , accessed on 23 May 2024). The database is called the Phytoplasma Disease and Symptom Database ( i PhyDSDB), which includes 1264 links to symptomatic images collected from 372 out of 945 plant speciesOpen asset ↗lines:30-40
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published23 Aug 2024AgronomyCited by 34 · OpenAlex ↗

Deep Learning-Based Methods for Multi-Class Rice Disease Detection Using Plant Images

RiceLeafClassificationObject detectionStress / disease detectionDisease symptoms / severityYield / yield components

Rapid and accurate diagnosis of rice diseases can prevent large-scale outbreaks and reduce pesticide overuse, thereby ensuring rice yield and quality. Existing research typically focuses on a limited number of rice diseases, which makes these studies less applicable to the diverse range of diseases currently affecting rice. Consequently, these studies fail to meet the detection needs of agricultural workers. Additionally, the lack of discussion regarding advanced detection algorithms in current research makes it difficult to determine the optimal application solution. To address these limitations, this study constructs a multi-class rice disease dataset comprising eleven rice diseases and one healthy leaf class. The resulting model is more widely applicable to a variety of diseases. Additionally, we evaluated advanced detection networks and found that DenseNet emerged as the best-performing model with an accuracy of 95.7%, precision of 95.3%, recall of 94.8%, F1 score of 95.0%, and a parameter count of only 6.97 M. Considering the current interest in transfer learning, this study introduced pre-trained weights from the large-scale, multi-class ImageNet dataset into the experiments. Among the tested models, RegNet achieved the best comprehensive performance, with an accuracy of 96.8%, precision of 96.2%, recall of 95.9%, F1 score of 96.0%, and a parameter count of only 3.91 M. Based on the transfer learning-based RegNet model, we developed a rice disease identification app that provides a simple and efficient diagnosis of rice diseases.

Why it matches plant phenotyping methodsイネ葉画像から病害状態を推定する深層学習手法を開発・比較し、データセットと診断アプリまで構築しており、植物病害表現型の取得・抽出が中心である。

abstractthis study constructs a multi-class rice disease dataset comprising eleven rice diseases and one healthy leaf class.
Reproduction assets foundThe paper's rice disease image dataset was partly acquired from a public Kaggle dataset (trumanrase/rice-leaf-diseases), which is a paper-specific, publicly available plant image asset used directly for the disease classification experiments. No author analysis code, trained model checkpoints, or supplementary deposits
Dataset · publicease dataset in- cludes 11 categories of rice diseases and 1 category of healthy leaves, totaling 11,281 images. The categories in the dataset are illustrated in Figure 1, and the number of images per cate- gory is detailed in Table 1. The dataset is divided into training, validation, and testing sets, with a ratio of 60:20:20 (https://www.kaggle.com/datasets/trumanrase/rice-leaf-diseases, accessed on 22 July 2024). The search engine method involves automatically downloading images by inputting keywords into Google using a Python script. The downloaded images are then filtered and cleaned to ensure data accuracy. The images for the disease categories bacterial leaf streak, Hispa, and rice shOpen asset ↗Kaggle · trumanrase/rice-leaf-diseasespdf-raw-page:3 lines:1-44
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published22 Aug 2024Data in briefCited by 1 · OpenAlex ↗

Tolerance to spittlebugs ( Aeneolamia varia ) in Urochloa spp. and Megathyrsus maximus grasses: A dataset for plant damage phenotyping.

Whole plant / canopy / plot / fieldStress / disease detectionStress response / tolerance

This dataset results from controlled experiments that assess the tolerance of Urochloa spp. and Megathyrsus maximus grasses to nymphal and adult spittlebug damage, particularly from Aeneolamia varia , which significantly impacts forage production in Neotropical regions. Data were collected under standardized conditions using high-throughput phenotyping methods, integrating image-capture techniques and analyses to ensure precise and consistent data acquisition. The dataset serves as a foundational resource for developing and validating computer vision models aimed at automated phenotyping, enabling accurate and high-throughput assessment of plant tolerance to spittlebug damage. Researchers can use the dataset to benchmark and compare different methodologies for plant damage assessment, fostering standardization and reproducibility in phenotyping studies.

Why it matches plant phenotyping methods植物の害虫被害耐性を画像ベースで高スループット評価するデータセットであり、コンピュータビジョンモデルの開発・検証と手法比較を目的とするため、フェノタイピング手法が中心です。

titleA dataset for plant damage phenotyping.
Reproduction assets foundThe paper is a Data in Brief article describing a public Dataverse deposit of 8318 high-resolution plant images and metadata for spittlebug damage phenotyping, with a direct URL matching an allowed URL.
Dataset · publicon Alliance of Bioversity International and CIAT in Palmira, Colombia (3°30′03.1″ N, 76°21′25.4″ W). Data accessibility Repository name: “Dataset: Tolerance to spittlebugs (Hemiptera: Cercopidae) in Urochloa spp. and Megathyrsus maximus grasses”. Data identification number: https://doi.org/10.7910/DVN/EGUVHA Direct URL to data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/EGUVHA Instructions for accessing these data: Access the public available dataset URL, download the files, and follow the instructions in the README file to decompress the dataset and preserve the intended structure of folders and files. 1. Value of the Data • High-throughput phenotyping: The datOpen asset ↗Dataverse · doi:10.7910/DVN/EGUVHAlines:1-50
Code / dataset availability confirmedCrossref · checked 14 Sept 2026
Published14 Aug 2024Copernicus GmbHCited by 1 · OpenAlex ↗

Mapping global leaf inclination angle (LIA) based on field measurement data

Field / plotLeafMorphology / geometry measurementArchitecture / morphology / geometry

Abstract. Leaf inclination angle (LIA), the angle between leaf surface normal and zenith directions, is a vital parameter in radiative transfer, rainfall interception, evapotranspiration, photosynthesis, and hydrological processes. Due to the difficulty in obtaining large-scale field measurement data, LIA is typically assumed to follow the spherical leaf distribution or simply considered constant for different plant types. However, the appropriateness of these simplifications and the global LIA distribution are still unknown. This study compiled global LIA measurements and generated the first global 500 m mean LIA (MLA) product by gap-filling the LIA measurement data using a random forest regressor. Different generation strategies were employed for noncrops and crops. The MLA product was evaluated by validating the nadir leaf projection function (G(0)) derived from the MLA product with high-resolution reference data. The global MLA is 41.47°±9.55°, and the value increases with latitude. The MLAs for different vegetation types follow the order of cereal crops (54.65°) > broadleaf crops (52.35°) > deciduous needleleaf forest (50.05°) > shrubland (49.23°) > evergreen needleleaf forest (47.13°) ≈ grassland (47.12°) > deciduous broadleaf forest (41.23°) > evergreen broadleaf forest (34.40°). Cross-validation shows that the predicted MLA presents a medium consistency (r = 0.75, RMSE = 7.15°) with the validation samples for noncrops, whereas crops show relatively lower correspondence (r = 0.48 and 0.60 for broadleaf crops and cereal crops) because of limited LIA measurements and strong seasonality. The global G(0) distribution is opposite to that of the MLA and agrees moderately with the reference data (r = 0.62, RMSE = 0.15). This study shows that the common spherical and constant LIA assumptions may underestimate the intercept capability for most vegetation. The MLA and G(0) products derived in this study would enhance our knowledge about global LIA and should greatly facilitate remote sensing retrieval and land surface modeling studies. The global MLA and G(0) products can be accessed at: Li, S. and Fang, H. 2024, https://doi.org/10.5281/zenodo.10940673.

Why it matches plant phenotyping methods全球の葉傾斜角という明示的な植物形質について、測定データの補間によるプロダクト作成と独立データによる検証が中心であり、単なる生態・農業研究での routine 測定ではない。

abstractThis study compiled global LIA measurements and generated the first global 500 m mean LIA (MLA) product by gap-filling the LIA measurement data using a random forest regressor.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicThe global MLA and G(0) products can be accessed at https://doi.org/10.5281/zenodo.12739662 (Li and Fang, 2025).Open asset ↗Zenodo · 10.5281/zenodo.12739662pdf-page:1 lines:1-51
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 Aug 2024Plant phenomics (Washington, D.C.)Cited by 26 · OpenAlex ↗

A Multi-Modal Open Object Detection Model for Tomato Leaf Diseases with Strong Generalization Performance Using PDC-VLD.

TomatoMultimodalLeafObject detectionDisease symptoms / severity

Precise disease detection is crucial in modern precision agriculture, especially in ensuring the health of tomato crops and enhancing agricultural productivity and product quality. Although most existing disease detection methods have helped growers identify tomato leaf diseases to some extent, these methods typically target fixed categories. When faced with new diseases, extensive and costly manual annotation is required to retrain the dataset. To overcome these limitations, this study proposes a multimodal model PDC-VLD based on the open-vocabulary object detection (OVD) technology within the VLDet framework, which can accurately identify new tomato leaf diseases without manual annotation by using only image-text pairs. First, we developed a progressive visual transformer-convolutional pyramid module (PVT-C) that effectively extracts tomato leaf disease features and optimizes anchor box positioning using the self-supervised learning algorithm DINO, suppressing interference from irrelevant backgrounds. Then, a context feature guided module (CFG) was adopted to address the low adaptability and recognition accuracy of the model in data-scarce environments. To validate the model's effectiveness, we constructed a tomato leaf disease image dataset containing 4 base classes and 2 new categories. Experimental results show that the PDC-VLD model achieved 61.2% on the main evaluation metric mAPnovel50 , and 56.4% on mAPnovel75 , 87.7% on mAPbase50 , 81.0% on mAPall50 , and 45.5% on average recall, outperforming existing OVD models. Our research provides an innovative solution for efficiently and accurately detecting new diseases, substantially reducing the need for manual annotation, and offering critical technical support and practical reference for agricultural workers.

Why it matches plant phenotyping methodsトマト葉の病害状態を画像から検出するマルチモーダルモデルを開発し、データセット構築と性能評価まで行っており、植物フェノタイピング手法が研究の中心である。

abstractthis study proposes a multimodal model PDC-VLD based on the open-vocabulary object detection (OVD) technology
Reproduction assets foundThe paper's Data Availability statement says all datasets used were uploaded to the authors' public GitHub repository (PDC-VLD), and the paper also uses the public Kaggle PlantVillage dataset as a plant image source. Bespoke portions of the dataset require contacting the corresponding author.
Dataset · publicAll datasets that were used and analyzed in this study have been uploaded to the website https://github.com/ZhouGuoXiong/PDC-VLD . Furthermore, for access to all bespoke datasets used in this study (comprising a total of 6,923 images and 13,864 texts), please contact the corresponding author.Open asset ↗ZhouGuoXiong/PDC-VLDlines:797-797
Code / dataset availability confirmedOpenAlex · checked 7 Sept 2026
Published12 Aug 2024AgricultureCited by 6 · OpenAlex ↗

SPCN: An Innovative Soybean Pod Counting Network Based on HDC Strategy and Attention Mechanism

SoybeanFruitCountingFruit / seed / panicle traits

Soybean pod count is a crucial aspect of soybean plant phenotyping, offering valuable reference information for breeding and planting management. Traditional manual counting methods are not only costly but also prone to errors. Existing detection-based soybean pod counting methods face challenges due to the crowded and uneven distribution of soybean pods on the plants. To tackle this issue, we propose a Soybean Pod Counting Network (SPCN) for accurate soybean pod counting. SPCN is a density map-based architecture based on Hybrid Dilated Convolution (HDC) strategy and attention mechanism for feature extraction, using the Unbalanced Optimal Transport (UOT) loss function for supervising density map generation. Additionally, we introduce a new diverse dataset, BeanCount-1500, comprising of 24,684 images of 316 soybean varieties with various backgrounds and lighting conditions. Extensive experiments on BeanCount-1500 demonstrate the advantages of SPCN in soybean pod counting with an Mean Absolute Error(MAE) and an Mean Squared Error(MSE) of 4.37 and 6.45, respectively, significantly outperforming the current competing method by a substantial margin. Its excellent performance on the Renshou2021 dataset further confirms its outstanding generalization potential. Overall, the proposed method can provide technical support for intelligent breeding and planting management of soybean, promoting the digital and precise management of agriculture in general.

Why it matches plant phenotyping methods大豆莢数という植物形質を画像から自動推定する手法を開発し、データセット上で性能評価しているため、植物フェノタイピング手法が研究の中心です。

abstractSoybean pod count is a crucial aspect of soybean plant phenotyping
Reproduction assets foundThe paper's SPCN analysis code is stated to be publicly available on GitHub. The BeanCount-1500 dataset itself has no public deposit or availability statement (authors must be contacted), so it does not qualify as a public asset.
Code · publicThe source code is available at https://github.com/johnhamtom/ soybean_counting_SPCN (accessed on 10 August 2024).Open asset ↗pdf-page:2 lines:1-58
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published10 Aug 2024Data in briefCited by 2 · OpenAlex ↗

Annotated image dataset of fire blight symptoms for object detection in orchards.

AppleField / plotRGB / grayscaleFlowerLeafStem / branchObject detectionDisease symptoms / severity

The monitoring of plant diseases in nurseries, breeding farms and orchards is essential for maintaining plant health. Fire blight ( Erwinia amylovora ) is still one of the most dangerous diseases in fruit production, as it can spread epidemically and cause enormous economic damage. All measures are therefore aimed at preventing the spread of the pathogen in the orchard and containing an infection at an early stage [1-6]. Efficiency in plant disease control benefits from the development of a digital monitoring system if the spatial and temporal resolution of disease monitoring in orchards can be increased [7]. In this context, a digital disease monitoring system for fire blight based on RGB images was developed for orchards. Between 2021 and 2024, data was collected on nine dates under different weather conditions and with different cameras. The data source locations in Germany were the experimental orchard of the Julius Kühn Institute (JKI), Institute of Plant Protection in Fruit Crops and Viticulture in Dossenheim, the experimental greenhouse of the Julius Kühn Institute for Resistance Research and Stress Tolerance in Quedlinburg and the experimental orchard of the JKI for Breeding Research on Fruit Crops located in Dresden-Pillnitz. The RGB images were taken on different apple genotypes after artificial inoculation with Erwinia amylovora , including cultivars, wild species and progeny from breeding. The presented ERWIAM dataset contains manually labelled RGB images with a size of 1280 × 1280 pixels of fire blight infected shoots, flowers and leaves in different stages of development as well as background images without symptoms. In addition, symptoms of other plant diseases were acquired and integrated into the ERWIAM dataset as a separate class. Each fire blight symptom was annotated with the Computer Vision Annotation Tool (CVAT [8]) using 2-point annotations (bounding boxes) and presented in YOLO 1.1 format (.txt files). The dataset contains a total of 1611 annotated images and 87 background images. This dataset can be used as a resource for researchers and developers working on digital systems for plant disease monitoring.

Why it matches plant phenotyping methodsRGB画像から植物病徴を検出するための注釈付きデータセットを開発・提示しており、植物病害状態の画像ベース表現型評価が中心である。

abstracta digital disease monitoring system for fire blight based on RGB images was developed for orchards.
Reproduction assets foundThe paper is a data descriptor for the ERWIAM dataset of annotated RGB images of fire blight symptoms, publicly deposited on Mendeley Data with a direct URL and DOI given in the text.
Dataset · publicTolerance located in Quedlinburg (Germany) [51°46ʹ22″N 11°08ʹ41″E] and at the experimental orchard of the JKI for Breeding Research on Fruit Crops located in Dresden-Pillnitz (Germany) [51°00ʹ01″N 13°53ʹ12"E]. Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/fpmnncmg84.1 Direct URL to data: https://data.mendeley.com/datasets/fpmnncmg84/1 1 Value of the Data • In the experimental greenhouse of the JKI-Quedlinburg Institute, around 2000 different genotypes of apple breeding material were artificially inoculated with Erwinia amylovora in 2021 and 2022, which could be used to record fire blight symptoms. The JKI-Dossenheim Institute has a heterogeneous appleOpen asset ↗Mendeley Data · 10.17632/fpmnncmg84.1lines:42-82
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published10 Aug 2024Data in briefCited by 1 · OpenAlex ↗

Dataset of Virginia Flue-cured Tobacco Leaf images based on stalk leaf position for classification tasks: A case of Tanzania.

TobaccoField / plotLeafClassification

Nicotiana tabacum is a kind of plant cultivated for its leaves used for manufacturing medicine and cigarettes. With the common name, the Tobacco plant is grown in many countries including China, Indonesia, Malawi and Tanzania just to mention a few. Literatures suggest a technical gap in the proper identification of grade labels for various parts of the plant. In addition, manual grading has resulted in various gaps and biases. To mitigate this, a data-driven grading solution is necessary. However, relevant datasets to train grade classifiers from various countries become of the essence. This article presents images concentrated on tobacco leaf plant position namely Leaf position which normally carries 23 grade labels. Due to high rainfall which swiped away the applied fertilizer on the tobacco plants in the farms, we failed to get images of one grade. Therefore, this research could capture and label 22 grade labels. Images of tobacco leaves based on the tobacco plant position were collected in Tanzania through participatory community research. Canon 5D mark III cameras with 100 mm micro lens were used to take pictures of tobacco leaves based on the tobacco plant position. Domain experts were used for image labelling and cleaning according to tobacco grade labels identified in Tanzania. The dataset carries 49,779 images, which can be used to develop machine learning models for tobacco leaf grade label identification. The collected dataset can be used to train models and enhance the performance of pre-trained models in any country of interest.

Why it matches plant phenotyping methodsタバコ葉の位置・等級を画像として収集し、専門家ラベル付きデータセットを構築しており、植物器官の状態・品質を画像から分類する再利用可能な方法資源が中心である。

abstractThis article presents images concentrated on tobacco leaf plant position namely Leaf position which normally carries 23 grade labels.
Reproduction assets foundThe paper is a data descriptor whose own tobacco leaf image dataset (49,779 images, 22 grade labels) is publicly deposited in the Harvard Dataverse with an explicit direct URL, making it a paper-specific, publicly actionable phenotyping image dataset.
Dataset · publicr tobacco leaves image sample, the names of each grade label within leaf position in the dataset were identified. Data source location Tanzania Tobacco Board (TTB), Tobacco Research Institute of Tanzania (TORITA) City/Town/Region: Tabora Country: Tanzania Data accessibility Repository name: Harvard Dataverse Direct URL to data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/TTPLFT 1 Value of the Data •Open asset ↗Harvard Dataverse · doi:10.7910/DVN/TTPLFTlines:1-49
Code / dataset availability confirmedOpenAlex · checked 7 Sept 2026
Published6 Aug 2024ForestsCited by 12 · OpenAlex ↗

YOLOTree-Individual Tree Spatial Positioning and Crown Volume Calculation Using UAV-RGB Imagery and LiDAR Data

Aerial / UAVLiDAR / point cloudRGB / grayscaleWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detectionSegmentationArchitecture / morphology / geometry

Individual tree canopy extraction plays an important role in downstream studies such as plant phenotyping, panoptic segmentation and growth monitoring. Canopy volume calculation is an essential part of these studies. However, existing volume calculation methods based on LiDAR or based on UAV-RGB imagery cannot balance accuracy and real-time performance. Thus, we propose a two-step individual tree volumetric modeling method: first, we use RGB remote sensing images to obtain the crown volume information, and then we use spatially aligned point cloud data to obtain the height information to automate the calculation of the crown volume. After introducing the point cloud information, our method outperforms the RGB image-only based method in 62.5% of the volumetric accuracy. The AbsoluteError of tree crown volume is decreased by 8.304. Compared with the traditional 2.5D volume calculation method using cloud point data only, the proposed method is decreased by 93.306. Our method also achieves fast extraction of vegetation over a large area. Moreover, the proposed YOLOTree model is more comprehensive than the existing YOLO series in tree detection, with 0.81% improvement in precision, and ranks second in the whole series for mAP50-95 metrics. We sample and open-source the TreeLD dataset to contribute to research migration.

Why it matches plant phenotyping methodsUAV-RGB画像とLiDARを用いて個体樹冠体積を推定する手法を開発・評価しており、単なる樹木位置検出を超えた植物形態形質の抽出が中心です。データセット公開も行っています。

abstractwe propose a two-step individual tree volumetric modeling method
Reproduction assets foundThe paper's authors explicitly state their analysis code (YOLOTree phenotyping/crown volume pipeline) is publicly available on GitHub, matching an allowed URL.
Code · publicOur code is available at: https://github.com/luotiger123/YOLOtree.Open asset ↗luotiger123/YOLOtreepdf-page:12 lines:1-67
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published2 Aug 2024Scientific reportsCited by 14 · OpenAlex ↗

The impact of fine-tuning paradigms on unknown plant diseases recognition.

ClassificationStress / disease detectionDisease symptoms / severity

Plant diseases pose significant threats to agriculture, impacting both food safety and public health. Traditional plant disease detection systems are typically limited to recognizing disease categories included in the training dataset, rendering them ineffective against new disease types. Although out-of-distribution (OOD) detection methods have been proposed to address this issue, the impact of fine-tuning paradigms on these methods has been overlooked. This paper focuses on studying the impact of fine-tuning paradigms on the performance of detecting unknown plant diseases. Currently, fine-tuning on visual tasks is mainly divided into visual-based models and visual-language-based models. We first discuss the limitations of large-scale visual language models in this task: textual prompts are difficult to design. To avoid the side effects of textual prompts, we futher explore the effectiveness of purely visual pre-trained models for OOD detection in plant disease tasks. Specifically, we employed five publicly accessible datasets to establish benchmarks for open-set recognition, OOD detection, and few-shot learning in plant disease recognition. Additionally, we comprehensively compared various OOD detection methods, fine-tuning paradigms, and factors affecting OOD detection performance, such as sample quantity. The results show that visual prompt tuning outperforms fully fine-tuning and linear probe tuning in out-of-distribution detection performance, especially in the few-shot scenarios. Notably, the max-logit-based on visual prompt tuning achieves an AUROC score of 94.8 % in the 8-shot setting, which is nearly comparable to the method of fully fine-tuning on the full dataset (95.2 % ), which implies that an appropriate fine-tuning paradigm can directly improve OOD detection performance. Finally, we visualized the prediction distributions of different OOD detection methods and discussed the selection of thresholds. Overall, this work lays the foundation for unknown plant disease recognition, providing strong support for the security and reliability of plant disease recognition systems. We will release our code at https://github.com/JiuqingDong/PDOOD to further advance this field.

Why it matches plant phenotyping methods植物画像から未知の病害状態を認識するOOD検出手法を、複数データセットでベンチマーク・比較しており、病害表現型の取得・推定手法が研究の中心です。

abstractwe employed five publicly accessible datasets to establish benchmarks for open-set recognition, OOD detection, and few-shot learning in plant disease recognition.
Reproduction assets foundThe paper's authors state they will release their analysis code (fine-tuning paradigms, OOD detection benchmarks, and detailed results) at a public GitHub repository, and note that more detailed results are available in that code repository. The datasets used (Cotton, Mango, Strawberry, Tomato, Plant Village) are cited
Code · publicWe will release our code at https://​github.​com/​Jiuqi​ngDong/​PDOOD to further advance this field.Open asset ↗Jiuqi​ngDong/​PDOODpdf-page:1 lines:1-62
Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Published31 Jul 2024arXivCited by 0 · OpenAlex ↗

High-throughput 3D shape completion of potato tubers on a harvester

PotatoField / plotLaboratory / benchtopLiDAR / point cloudRGB-D / ToFFruitRootWhole plant / canopy / plot / field2D/3D reconstructionYield / biomass estimation

Potato yield is an important metric for farmers to further optimize their cultivation practices. Potato yield can be estimated on a harvester using an RGB-D camera that can estimate the three-dimensional (3D) volume of individual potato tubers. A challenge, however, is that the 3D shape derived from RGB-D images is only partially completed, underestimating the actual volume. To address this issue, we developed a 3D shape completion network, called CoRe++, which can complete the 3D shape from RGB-D images. CoRe++ is a deep learning network that consists of a convolutional encoder and a decoder. The encoder compresses RGB-D images into latent vectors that are used by the decoder to complete the 3D shape using the deep signed distance field network (DeepSDF). To evaluate our CoRe++ network, we collected partial and complete 3D point clouds of 339 potato tubers on an operational harvester in Japan. On the 1425 RGB-D images in the test set (representing 51 unique potato tubers), our network achieved a completion accuracy of 2.8 mm on average. For volumetric estimation, the root mean squared error (RMSE) was 22.6 ml, and this was better than the RMSE of the linear regression (31.1 ml) and the base model (36.9 ml). We found that the RMSE can be further reduced to 18.2 ml when performing the 3D shape completion in the center of the RGB-D image. With an average 3D shape completion time of 10 milliseconds per tuber, we can conclude that CoRe++ is both fast and accurate enough to be implemented on an operational harvester for high-throughput potato yield estimation. CoRe++'s high-throughput and accurate processing allows it to be applied to other tuber, fruit and vegetable crops, thereby enabling versatile, accurate and real-time yield monitoring in precision agriculture. Our code, network weights and dataset are publicly available at https://github.com/UTokyo-FieldPhenomics-Lab/corepp.git.

Why it matches plant phenotyping methodsRGB-D画像からジャガイモ塊茎の3D形状を補完し、体積・収量を推定する手法を開発・検証しており、植物フェノタイピング手法が研究の中心である。

abstractwe developed a 3D shape completion network, called CoRe++, which can complete the 3D shape from RGB-D images.
Reproduction assets foundThe paper's abstract explicitly states that the authors' code, network weights, and the potato tuber RGB-D/3D point cloud dataset are publicly available at the authors' GitHub repository (UTokyo-FieldPhenomics-Lab/corepp), which is a paper-specific, public, actionable asset for the CoRe++ phenotyping analysis.
Code · publicOur code, network weights and dataset are publicly available at https://github.com/UTokyo-FieldPhenomics-Lab/corepp.git .Open asset ↗UTokyo-FieldPhenomics-Lab/corepplines:1-93
Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Published30 Jul 2024arXivCited by 4 · OpenAlex ↗

PLANesT-3D: A new annotated dataset for segmentation of 3D plant point clouds

Pepper / chilliPhotogrammetry / SfM / MVSLiDAR / point cloudLeafStem / branchSegmentation

Creation of new annotated public datasets is crucial in helping advances in 3D computer vision and machine learning meet their full potential for automatic interpretation of 3D plant models. Despite the proliferation of deep neural network architectures for segmentation and phenotyping of 3D plant models in the last decade, the amount of data, and diversity in terms of species and data acquisition modalities are far from sufficient for evaluation of such tools for their generalization ability. To contribute to closing this gap, we introduce PLANesT-3D; a new annotated dataset of 3D color point clouds of plants. PLANesT-3D is composed of 34 point cloud models representing 34 real plants from three different plant species: \textit{Capsicum annuum}, \textit{Rosa kordana}, and \textit{Ribes rubrum}. Both semantic labels in terms of "leaf" and "stem", and organ instance labels were manually annotated for the full point clouds. PLANesT-3D introduces diversity to existing datasets by adding point clouds of two new species and providing 3D data acquired with the low-cost SfM/MVS technique as opposed to laser scanning or expensive setups. Point clouds reconstructed with SfM/MVS modality exhibit challenges such as missing data, variable density, and illumination variations. As an additional contribution, SP-LSCnet, a novel semantic segmentation method that is a combination of unsupervised superpoint extraction and a 3D point-based deep learning approach is introduced and evaluated on the new dataset. The advantages of SP-LSCnet over other deep learning methods are its modular structure and increased interpretability. Two existing deep neural network architectures, PointNet++ and RoseSegNet, were also tested on the point clouds of PLANesT-3D for semantic segmentation.

Why it matches plant phenotyping methods3D植物点群の注釈付きデータセットを構築し、植物器官のセマンティック・インスタンス分割手法を開発・評価しており、植物フェノタイピング手法が中心である。

abstractwe introduce PLANesT-3D; a new annotated dataset of 3D color point clouds of plants.
Reproduction assets foundThe paper introduces PLANesT-3D, an annotated 3D plant point cloud dataset, and SP-LSCnet segmentation code, both explicitly stated as publicly available at the authors' Aperta record and GitHub repository.
Dataset · publicThe PLANesT-3D dataset is publicly available at https://aperta.ulakbim.gov.tr/record/286354 and https://github.com/visionlab-ogu/PLANesT-3D/tree/main/dataOpen asset ↗aperta.ulakbim.gov.tr · 286354lines:83-145
Dataset · publicThe 2D color images for all the 34 plants together with their estimated camera poses and parameters are also open to the public to provide input data for recent 3D reconstruction techniques 3 3 3 The data is available at https://github.com/visionlab-ogu/PLANesT-3D/tree/main/data .Open asset ↗github.com/visionlab-ogu/PLANesT-3Dlines:494-505
Code · publicThe code for SP-LSCnet is available at https://github.com/visionlab-ogu/PLANesT-3DOpen asset ↗github.com/visionlab-ogu/PLANesT-3Dlines:146-154
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published24 Jul 2024Copernicus GmbHCited by 1 · OpenAlex ↗

Partitioning of water and CO 2 fluxes at NEON sites into soil and plant components: a five-year dataset for spatial and temporal analysis

Field / plotWhole plant / canopy / plot / fieldPhysiological trait estimationGrowth / time-series analysisPhotosynthesis / fluorescenceWater status / transpiration

Abstract. Long-term time series of transpiration, evaporation, plant photosynthesis, and soil respiration are essential for addressing numerous research questions related to ecosystem functioning. However, quantifying these fluxes is challenging due to the lack of reliable and direct measurement techniques, which has left gaps in the understanding of their temporal cycles and spatial variability. To help address this open challenge, we generated a dataset of these four components by implementing five (conventional and novel) approaches to partition total ET and CO2 fluxes into plant and soil fluxes across 47 NEON sites. The final dataset (https://doi.org/10.5281/zenodo.12191876) spans a five-year period and covers various ecosystems, including forests, grasslands, and agricultural terrain. This is the first comprehensive dataset covering such a wide spatial and temporal distribution. Overall, we observed good agreement across most methods for ET components, increasing the reliability of these estimates. Partitioning of CO2 components was found to be less robust and more dependent on prior knowledge of water-use efficiency. This dataset has several potential future applications, such as addressing critical questions regarding the response of ecosystems to extreme weather events, which are expected to become more severe and frequent with climate change.

Why it matches plant phenotyping methods植物・土壌フラックスを分離推定する5手法を47地点で実装し、手法間の一致度を評価した長期データセットであり、植物の生理状態(蒸散・光合成)の取得・推定法が中心的です。

abstractwe generated a dataset of these four components by implementing five (conventional and novel) approaches to partition total ET and CO2 fluxes into plant and soil fluxes across 47 NEON sites.
Reproduction assets foundThe paper's five-year NEON flux-partitioning dataset and the authors' partitioning-method scripts are explicitly deposited on Zenodo with public DOIs.
Dataset · publicows the availability of flux components as a fraction of the total number of half- hour periods in the record. Overall, all the methods cover a similar temporal distribution of flux partitioning and are potential candidates for ensemble averaging. 4 Description of the final dataset The final dataset is available for download at https://doi.org/10.5281/zenodo.12191876 (Zahn and Bou- Zeid, 2024). It is organized into different folders for each site, with each site containing a .csv file for each method. This format is selected to be user-friendly and accessible in various programming languages and software packages. For FVS and CECw, in addition to their ensemble averages for https://doi.org/Open asset ↗Zenodo · 10.5281/zenodo.12191876pdf-raw-page:9 lines:136-149
Code · publicnthesis, transpiration and stomatal conduc- tance: potential and limitations, Plant Cell Environ., 35, 657– 667, https://doi.org/10.1111/j.1365-3040.2011.02451.x, 2011. Zahn, E.: einaraz/PartitioningMethods: Processing Eddy- Covariance Data: Five Evapotranspiration Flux Parti- tioning Methods (v1.0.1) [Software], Zenodo [code], https://doi.org/10.5281/zenodo.11510363, 2024. Zahn, E. and Bou-Zeid, E.: Partitioning of water and CO2 fluxes at NEON sites into soil and plant components: a five-year dataset for spatial and temporal analysis [dataset], Zenodo [data set], https://doi.org/10.5281/zenodo.12191876, 2024. Zahn, E., Chor, T. L., and Dias, N. L.: A Simple Methodology for Quality ControlOpen asset ↗Zenodo · 10.5281/zenodo.11510363pdf-raw-page:22 lines:1-58
Code / dataset availability confirmedbioRxiv · checked 14 Sept 2026
Published23 Jul 2024bioRxivCited by 0 · OpenAlex ↗

Off-the-shelf image analysis models outperform human visual assessment in identifying genes controlling seed color variation in sorghum

Seed / grainMorphology / geometry measurementPigment / colour / senescence

Seed color is a complex phenotype linked to both the impact of grains on human health and consumer acceptance of new crop varieties. Today seed color is often quantified via either qualitative human assessment or biochemical assays for specific colored metabolites. Imaging-based approaches have the potential to be more quantitative than human scoring while lower cost than biochemical assays. We assessed the feasibility of employing image analysis tools trained on rice (Oryza sativa) or wheat (Triticum aestivum) seeds to quantify seed color in sorghum (Sorghum bicolor ) using a dataset of > 1,500 images. Quantitative measurements of seed color from images were substantially more consistent across biological replicates than human assessment. Genome-wide association studies conducted using color phenotypes for 682 sorghum genotypes identified more signals near known seed color genes in sorghum with stronger support than manually scored seed color for the same experiment. Previously unreported genomic intervals linked to variation in seed color in our study co-localized with a gene encoding an enzyme in the biosynthetic pathway leading to anthocyanins, tannins, and phlobaphenes - colored metabolites in sorghum seeds - and with the sorghum ortholog of a transcription factor shown to regulate several enzymes in the same pathway in rice. The cross-species transferability of image analysis tools, without the retraining, may aid efforts to develop higher value and health-promoting crop varieties in sorghum and other specialty and orphan grain crops.

Why it matches plant phenotyping methods画像解析モデルを用いたソルガム種子色の定量化を中心に、手作業評価との再現性比較と遺伝解析による妥当性評価を行っているため、植物表現型計測手法として採用。

abstractWe assessed the feasibility of employing image analysis tools trained on rice (Oryza sativa) or wheat (Triticum aestivum) seeds to quantify seed color in sorghum (Sorghum bicolor ) using a dataset of > 1,500 images.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · public12 using codes available in https://github.com/NikeeShrestha/SorghumSeedSegmentation.Open asset ↗NikeeShrestha/SorghumSeedSegmentationpdf-page:6 lines:1-60
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published20 Jul 2024Data in briefCited by 10 · OpenAlex ↗

A novel groundnut leaf dataset for detection and classification of groundnut leaf diseases.

Peanut / groundnutField / plotRGB / grayscaleLeafClassificationStress / disease detectionDisease symptoms / severity

Groundnut (Arachis hypogaea) is a widely cultivated legume crop that plays a vital role in global agriculture and food security. It is a major source of vegetable oil and protein for human consumption, as well as a cash crop for farmers in many regions. Despite the importance of this crop to household food security and income, diseases, particularly Leaf spot (early and late), Alternaria leaf spot, Rust, and Rosette, have had a significant impact on its production. Deep learning (DL) techniques, especially convolutional neural networks (CNNs), have demonstrated significant ability for early diagnosis of the plant leaf diseases. However, the availability of groundnut-specific datasets for training and evaluation of DL models is limited, hindering the development and benchmarking of groundnut-related deep learning applications. Therefore, this study provides a dataset of groundnut leaf images, both diseased and healthy, captured in real cultivation fields at Ramchandrapur, Purba Medinipur, West Bengal, using a smartphone camera. The dataset contains a total of 1720 original images, that can be utilized to train DL models to detect groundnut leaf diseases at an early stage. Additionally, we provide baseline results of applying state-of-the-art CNN architectures on the dataset for groundnut disease classification, demonstrating the potential of the dataset for advancing groundnut-related research using deep learning. The aim of creating this dataset is to facilitate in the creation of sophisticated methods that will aid farmers accurately identify diseases and enhance groundnut yields.

Why it matches plant phenotyping methods落花生葉の健全・病葉画像データセットを提供し、植物病害状態の画像ベース判定を可能にすることが中心で、ベースライン評価も含むため。

abstractTherefore, this study provides a dataset of groundnut leaf images, both diseased and healthy, captured in real cultivation fields
Reproduction assets foundThe paper's own groundnut leaf image dataset (1720 images, diseased and healthy) is publicly deposited on Mendeley Data with an explicit DOI and direct URL, matching an allowed URL.
Dataset · publiccategorised based on disease criteria with the assistance of a pathologist. Data source location Ramchandrapur, Purba Medinipur, West Bengal, India, Pin: 721429 Latitude 21.930146 and Longitude 87.556852 Data accessibility Repository name: Mendeley Data. Data identification number: DOI: 10.17632/x6x5jkk873.2 Direct URL to data: https://data.mendeley.com/datasets/x6x5jkk873/2 Instructions for accessing these data: All the image can be downloaded by the following link: https://data.mendeley.com/datasets/x6x5jkk873/2 1. Value of the Data • We address four prominent diseases that specifically target groundnut leaves, causing significant damage to numerous groundnut fields. Researchers and practiOpen asset ↗Mendeley Data · 10.17632/x6x5jkk873.2lines:1-47
Code / dataset availability confirmedarXiv · checked 14 Sept 2026
Published18 Jul 2024arXivCited by 0 · OpenAlex ↗

A Dataset and Benchmark for Shape Completion of Fruits for Agricultural Robotics

Pepper / chilliGreenhouseLaboratory / benchtopRGB-D / ToFFruit2D/3D reconstruction

As the world population is expected to reach 10 billion by 2050, our agricultural production system needs to double its productivity despite a decline of human workforce in the agricultural sector. Autonomous robotic systems are one promising pathway to increase productivity by taking over labor-intensive manual tasks like fruit picking. To be effective, such systems need to monitor and interact with plants and fruits precisely, which is challenging due to the cluttered nature of agricultural environments causing, for example, strong occlusions. Thus, being able to estimate the complete 3D shapes of objects in presence of occlusions is crucial for automating operations such as fruit harvesting. In this paper, we propose the first publicly available 3D shape completion dataset for agricultural vision systems. We provide an RGB-D dataset for estimating the 3D shape of fruits. Specifically, our dataset contains RGB-D frames of single sweet peppers in lab conditions but also in a commercial greenhouse. For each fruit, we additionally collected high-precision point clouds that we use as ground truth. For acquiring the ground truth shape, we developed a measuring process that allows us to record data of real sweet pepper plants, both in the lab and in the greenhouse with high precision, and determine the shape of the sensed fruits. We release our dataset, consisting of almost 7,000 RGB-D frames belonging to more than 100 different fruits. We provide segmented RGB-D frames, with camera intrinsics to easily obtain colored point clouds, together with the corresponding high-precision, occlusion-free point clouds obtained with a high-precision laser scanner. We additionally enable evaluation of shape completion approaches on a hidden test set through a public challenge on a benchmark server.

Why it matches plant phenotyping methods果実の3D形状という植物器官の形態形質を対象に、RGB-D画像・高精度点群・評価用ベンチマークを構築しており、形状取得と推定の方法論が中心である。

abstractWe provide an RGB-D dataset for estimating the 3D shape of fruits.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicOur development toolkit including a data loader is available at: https://github.com/PRBonn/shape_completion_toolkit for handling the dataset and computing metrics.Open asset ↗PRBonn/shape_completion_toolkitlines:55-81
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published17 Jul 2024Data in briefCited by 4 · OpenAlex ↗

Fruit surface temperature data at different ripeness stages and ambient temperature provided as temperature-annotated 3D point clouds of apple trees.

AppleField / plotLiDAR / point cloudThermalFruitWhole plant / canopy / plot / fieldSegmentationFruit / seed / panicle traitsPlant / canopy temperature

By means of a unique, low vibration circular conveyor system, plant sensors capturing light detection and ranging (LiDAR) unit and thermal camera were moved on the same route, around blocks of apple trees, with seven Malus x domestica Borkh. 'Gala' apple trees in each block. Measurements took place four times during the season. Additionally at harvest, diurnal courses were recorded with 18 readings during three days. The data are provided as [i] raw data (3D point clouds of 3 blocks of trees scanned from right and left sides and thermal images), [ii] processed 3D point clouds of canopies annotated with temperature data from the thermal camera, and [iii] manually segmented 3D point clouds of fruit, representing the spatially-resolved fruit surface temperature (FST). Manual FST readings are provided on each measuring date and during diurnal courses. The fruit data are capturing 1236 FST, providing temperature distribution as 3D point cloud and one manually recorded reference FST per fruit. Additionally, fruit size and colour were measured for each fruit, despite for the first date, when fruit were too small for colour readings. Weather data are provided from a station located in the orchard. Usage of data could be (a) in developing methodology for 3D point cloud processing based on raw data, accomplished with reference FST data. Furthermore, (b) the pre-processed point clouds of fruit surface temperature can be reused in ecophysiological studies related to global warming, optimizing fruit production systems, and other. Because the sensors and trees were measured from the same angle and distance, time series analysis of the canopies would be possible.

Why it matches plant phenotyping methodsLiDARと熱画像を統合し、果実表面温度を3D点群として取得・注釈化した再利用可能なデータセットであり、植物表現型取得手法とデータ提供が中心です。

abstractplant sensors capturing light detection and ranging (LiDAR) unit and thermal camera were moved on the same route, around blocks of apple trees
Reproduction assets foundThe paper is a Data in Brief article describing a public Zenodo deposit containing the paper's own phenotyping measurements: raw LiDAR point clouds, thermal images, temperature-annotated 3D point clouds of apple canopies, 1236 manually segmented fruit point clouds with FST reference readings, fruit size/colour data, an
Dataset · publicocation The conveyor system is located 52.4673340479, 12.9606589643 in the experimental station of Leibniz Institute for Agricultural Engineering and Bioeconomy in Potsdam, Germany (ATB). Data repository is stored on Zenodo server [ 1 ] Data accessibility Repository name: Zenodo Doi: https://doi.org/10.5281/zenodo.10792723 url: https://zenodo.org/records/10792723 Related research article 1 Value of the Data • The stationary conveyor system enabled repeated readings of apple tree canopies, with minimum vibration due to electric engine of the conveyor, and equal geometry between sensors and samples in all measurements. The value of 3D point clouds obtained with LiDAR sensor was enhanced bOpen asset ↗Zenodo · 10.5281/zenodo.10792723lines:46-71
Code / dataset availability confirmedEurope PMC · checked 13 Sept 2026
Published10 Jul 2024Data in briefCited by 3 · OpenAlex ↗

PC4C_CAPSI: Image data of capsicum plant growth in protected horticulture.

Pepper / chilliGreenhouseLiDAR / point cloudRGB-D / ToFWhole plant / canopy / plot / fieldMorphology / geometry measurementOrgan identification2D/3D reconstructionSegmentationGrowth / development / phenology

Feeding the increasing global population and reducing the carbon footprint of agricultural activities are two critical challenges of our century. Growing crops under protected horticulture and precise crop monitoring have emerged to address these challenges. Crop monitoring in commercial protected facilities remains mostly manual and labour intensive. Using computer vision to solve specific problems in image-based crop monitoring in these compact and complex growth environments is currently hindered by the scarcity of available data. We collected an RGBD dataset for vertically supported, hydroponically-grown capsicum plants in a commercial-scale glasshouse facility to fill this gap. Data were collected weekly using a single top-angled stereo camera mounted on a mobile platform running between the hydroponic gutters. The RGBD streams covered 80 % of the crop growing season in three different light conditions. The metadata include camera configurations and light condition information. Manually measured plant heights of ten selected plants per gutter are provided as ground truth. The images covered the whole plants and focused on the top third. This dataset will support research on plant height estimation, plant organ identification, object segmentation, organ measurements, 3D reconstruction, 3D data processing, and depth noise reduction. The usability of the dataset has been successfully demonstrated in a previously published study on plant height estimation using machine learning and 3D point cloud.

Why it matches plant phenotyping methods植物の草丈推定や器官計測を目的としたRGBD画像データセットを構築し、地上真値も提供しているため、植物フェノタイピング手法・データセットが中心です。

abstractWe collected an RGBD dataset for vertically supported, hydroponically-grown capsicum plants in a commercial-scale glasshouse facility to fill this gap.
Reproduction assets foundThis data article directly deposits its paper-specific phenotyping assets: the PC4C_CAPSI RGBD image dataset (Rosbag streams, JSON metadata, manual plant-height ground truth) on the Western Sydney University ResearchDirect repository, and the authors' RGBD processing code (image extraction, depth correction, 3D reconss
Dataset · publicData accessibility Repository name: Image Data of Capsicum Plant Growth in Protected Horticulture: PC4C_CAPSI. [ 1 ] Data identification number: 10.26183/1A0R-E318 Direct URL to data: https://rds.westernsydney.edu.au/Institutes/HIE/2024/Jayasuriya_N/Open asset ↗rds.westernsydney.edu.aulines:1-40
Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Published8 Jul 2024arXiv

High-Throughput Phenotyping using Computer Vision and Machine Learning

PoplarField / plotLeafWhole plant / canopy / plot / fieldClassificationSegmentationStress / disease detectionLeaf traitsPigment / colour / senescenceStress response / tolerance

High-throughput phenotyping refers to the non-destructive and efficient evaluation of plant phenotypes. In recent years, it has been coupled with machine learning in order to improve the process of phenotyping plants by increasing efficiency in handling large datasets and developing methods for the extraction of specific traits. Previous studies have developed methods to advance these challenges through the application of deep neural networks in tandem with automated cameras; however, the datasets being studied often excluded physical labels. In this study, we used a dataset provided by Oak Ridge National Laboratory with 1,672 images of Populus Trichocarpa with white labels displaying treatment (control or drought), block, row, position, and genotype. Optical character recognition (OCR) was used to read these labels on the plants, image segmentation techniques in conjunction with machine learning algorithms were used for morphological classifications, machine learning models were used to predict treatment based on those classifications, and analyzed encoded EXIF tags were used for the purpose of finding leaf size and correlations between phenotypes. We found that our OCR model had an accuracy of 94.31% for non-null text extractions, allowing for the information to be accurately placed in a spreadsheet. Our classification models identified leaf shape, color, and level of brown splotches with an average accuracy of 62.82%, and plant treatment with an accuracy of 60.08%. Finally, we identified a few crucial pieces of information absent from the EXIF tags that prevented the assessment of the leaf size. There was also missing information that prevented the assessment of correlations between phenotypes and conditions. However, future studies could improve upon this to allow for the assessment of these features.

Why it matches plant phenotyping methods植物画像からラベル情報を読み取り、画像分割・機械学習で葉形、色、斑点などの形態形質を抽出・分類する手法が研究の中心であり、植物フェノタイピング手法の開発・適用に該当する。

abstractimage segmentation techniques in conjunction with machine learning algorithms were used for morphological classifications
Reproduction assets foundThe paper's authors explicitly state that all analysis code (OCR label reading, leaf segmentation, morphology classification, treatment prediction) is publicly available under the MIT License on their GitHub repository. The underlying ORNL image dataset is not stated to be publicly available, so only the code asset is.
Code · publicSince a pre-trained segmentation model (the SAM) was used in this study, researchers could attempt to build segmentation models fine-tuned to only recognize leaves, which could increase model efficiency and provide more consistent results. 6 Code Availability All code is publicly available under the MIT License on GitHub here: https://github.com/vivaansinghvi07/smoky-mountain-data-comp . Acknowledgements We thank Dr. Ty Frazier at Oak Ridge National Laboratory for his helpful suggestions and mentoring throughout this project. References Arya et al. (2022) Arya, S., Sandhu, K.S., Singh, J., Kumar, S., 2022. Deep learning: As the new frontier in high-throughput plant phenotyping. Euphytica 218Open asset ↗vivaansinghvi07/smoky-mountain-data-complines:272-401
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published8 Jul 2024HeliyonCited by 16 · OpenAlex ↗

Bacterial-fungicidal vine disease detection with proximal aerial images.

GrapevineAerial / UAVFruitLeafObject detectionStress / disease detectionDisease symptoms / severity

Vine disease detection is considered one of the most crucial components in precision viticulture. It serves as an input for several further modules, including mapping, automatic treatment, and spraying devices. In the last few years, several approaches have been proposed for detecting vine disease based on indoor laboratory conditions or large-scale satellite images integrated with machine learning tools. However, these methods have several limitations, including laboratory-specific conditions or limited visibility into plant-related diseases. To overcome these limitations, this work proposes a low-altitude drone flight approach through which a comprehensive dataset about various vine diseases from a large-scale European dataset is generated. The dataset contains typical diseases such as downy mildew or black rot affecting the large variety of grapes including Muscat of Hamburg, Alphonse Lavallée, Grasă de Cotnari, Rkatsiteli, Napoca, Pinot blanc, Pinot gris, Chambourcin, Fetească regală, Sauvignon blanc, Muscat Ottonel, Merlot, and Seyve-Villard 18402. The dataset contains 10,000 images and more than 100,000 annotated leaves, verified by viticulture specialists. Grape bunches are also annotated for yield estimation. Further, tests were made against state-of-the-art detection methods on this dataset, focusing also on viable solutions on embedded devices, including Android-based phones or Nvidia Jetson boards with GPU. The datasets, as well as the customized embedded models, are available on the project webpage.

Why it matches plant phenotyping methodsブドウ葉の病徴を低高度ドローン画像で検出する大規模データセットを構築し、葉アノテーションと手法比較・組込み機器での評価を行っており、植物状態の画像ベース計測が中心です。

abstractthis work proposes a low-altitude drone flight approach through which a comprehensive dataset about various vine diseases from a large-scale European dataset is generated.
Reproduction assets foundThe authors state their UAV vine-disease dataset (10,000 images, 100,000+ annotated leaves), preprocessing scripts, and pre-trained embedded models are publicly available on the project website, whose URL (github.com/tamaslevente/vineye) appears in the supplied blocks. Other URLs (ultralytics, CVAT, labelImg, Zenodo, k
Dataset · publicThe dataset and useful preprocessing scripts and pre-trained models are available on the project website.Open asset ↗lines:43-81
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published6 Jul 2024Data in briefCited by 21 · OpenAlex ↗

Multi-format open-source sweet orange leaf dataset for disease detection, classification, and analysis.

CitrusLeafClassificationStress / disease detectionDisease symptoms / severity

In Bangladesh, sweet orange cultivation has been popular among fruit growers as the fruit is in demand. However, the disease of sweet oranges decreases fruit production. Research suggests that computer-aided disease diagnosis and machine learning (IML) models can improve fruit production by detecting and classifying diseases. In this line, a dataset of sweet oranges is required to diagnose the disease. Moreover, like many other fruits, sweet orange disease may vary from country to country. Therefore, in Bangladesh, a sweet orange dataset is required. Lastly, since different ML algorithms require datasets in various formats, only a few existing datasets fulfil the necessity. To fulfil the limitations, a sweet orange dataset in Bangladesh is collected. The dataset was collected in August and comprises high-quality images documenting multiple disease conditions, including Citrus Canker, Citrus Greening, Citrus Mealybugs, Die Back, Foliage Damage, Spiny Whitefly, Powdery Mildew, Shot Hole, Yellow Dragon, Yellow Leaves, and Healthy Leaf . These images provide an opportunity to apply machine learning and computer vision techniques to detect and classify diseases. This dataset aims to help researchers advance agri engineering through ML. Other sweet orange growing countries with having similar environments may find helpful information. Lastly, such experiments using our dataset will assist farmers in taking preventive measures and minimising economic losses.

Why it matches plant phenotyping methodsスイートオレンジ葉の病害状態を画像で記録した再利用可能なデータセットの構築が中心であり、植物病害表現型の画像ベース推定に該当する。

abstracta dataset of sweet oranges is required to diagnose the disease
Reproduction assets foundThe paper is a Data in Brief article describing a sweet orange leaf disease image dataset (5,813 images, 11 classes, plus TXT annotations), publicly deposited on Mendeley Data with an explicit direct URL and DOI (10.17632/f7cr74mwpj.1). This is a paper-specific public plant image/annotation dataset directly reproducing
Dataset · publicl . Data source location City: Khemerdia, Bheramara, Kushtia Country: Bangladesh Latitude and longitude (and GPS coordinates, if possible) for collected samples/data: 35.3602°N and 113.9505°E, Altitude: 75 msl Data accessibility Repository name: Mendeley Data Data identification number: 10.17632/f7cr74mwpj.1 Direct URL to data: https://data.mendeley.com/datasets/f7cr74mwpj/1 1. Value of the Data •Open asset ↗Mendeley Data · 10.17632/f7cr74mwpj.1lines:1-46
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published2 Jul 2024Frontiers in plant scienceCited by 4 · OpenAlex ↗

CTHNet: a network for wheat ear counting with local-global features fusion based on hybrid architecture.

WheatRGB / grayscalePanicle / ear / spikeCountingFruit / seed / panicle traits

Accurate wheat ear counting is one of the key indicators for wheat phenotyping. Convolutional neural network (CNN) algorithms for counting wheat have evolved into sophisticated tools, however because of the limitations of sensory fields, CNN is unable to simulate global context information, which has an impact on counting performance. In this study, we present a hybrid attention network (CTHNet) for wheat ear counting from RGB images that combines local features and global context information. On the one hand, to extract multi-scale local features, a convolutional neural network is built using the Cross Stage Partial framework. On the other hand, to acquire better global context information, tokenized image patches from convolutional neural network feature maps are encoded as input sequences using Pyramid Pooling Transformer. Then, the feature fusion module merges the local features with the global context information to significantly enhance the feature representation. The Global Wheat Head Detection Dataset and Wheat Ear Detection Dataset are used to assess the proposed model. There were 3.40 and 5.21 average absolute errors, respectively. The performance of the proposed model was significantly better than previous studies.

Why it matches plant phenotyping methods小麦穂数という植物形質をRGB画像から推定する深層学習手法を開発し、複数データセットで性能評価しており、表現型取得・抽出法が中心である。

abstractAccurate wheat ear counting is one of the key indicators for wheat phenotyping.
Reproduction assets foundThe paper uses two publicly available wheat ear image datasets (GWHD and WEDD) as its phenotyping inputs, with explicit public URLs in the data availability statement. No authors' analysis code or trained model is deposited.
Dataset · publics generalization ability. This will provide real-time and accurate information for agricultural production, help farmers make scientific decisions, and improve crop management and yield. Data availability statement Publicly available datasets were analyzed in this study. This data can be found here: http://www.global-wheat.com/ https://github.com/simonMadec . Author contributions QH: Conceptualization, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing – review & editing. WL: Conceptualization, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. YZ: Software, Writing – review & editing. TR: SoftwaOpen asset ↗https://github.com/simonMadeclines:388-410
Code / dataset availability confirmedEurope PMC · checked 7 Sept 2026
Published21 Jun 2024Journal of imagingCited by 7 · OpenAlex ↗

Efficient Wheat Head Segmentation with Minimal Annotation: A Generative Approach.

WheatPanicle / ear / spikeSegmentation

Deep learning models have been used for a variety of image processing tasks. However, most of these models are developed through supervised learning approaches, which rely heavily on the availability of large-scale annotated datasets. Developing such datasets is tedious and expensive. In the absence of an annotated dataset, synthetic data can be used for model development; however, due to the substantial differences between simulated and real data, a phenomenon referred to as domain gap, the resulting models often underperform when applied to real data. In this research, we aim to address this challenge by first computationally simulating a large-scale annotated dataset and then using a generative adversarial network (GAN) to fill the gap between simulated and real images. This approach results in a synthetic dataset that can be effectively utilized to train a deep-learning model. Using this approach, we developed a realistic annotated synthetic dataset for wheat head segmentation. This dataset was then used to develop a deep-learning model for semantic segmentation. The resulting model achieved a Dice score of 83.4% on an internal dataset and Dice scores of 79.6% and 83.6% on two external datasets from the Global Wheat Head Detection datasets. While we proposed this approach in the context of wheat head segmentation, it can be generalized to other crop types or, more broadly, to images with dense, repeated patterns such as those found in cellular imagery.

Why it matches plant phenotyping methodsコムギ穂の画像セグメンテーション手法と合成データセットを開発し、内部・外部データセットで性能検証しており、植物フェノタイピング手法が中心である。

abstractwe developed a realistic annotated synthetic dataset for wheat head segmentation.
Reproduction assets foundThe paper's wheat head segmentation datasets (synthetic, GAN-generated, and evaluation sets) are publicly available at the authors' stated URL; the analysis code is only available on request.
Dataset · publicPublicly available datasets were utilized in this study. These data can be found here: https://www.cs.usask.ca/ftp/pub/whs/ (accessed on 1 June 2023). The code used to generate synthetic data presented in this study are available on request.Open asset ↗lines:74-261
Code / dataset availability confirmedOpenAlex · checked 7 Sept 2026
Published13 Jun 2024PLoS ONECited by 3 · OpenAlex ↗

TaeC: A manually annotated text dataset for trait and phenotype extraction and entity linking in wheat breeding literature

WheatAnnotation / quality controlClassification

Wheat varieties show a large diversity of traits and phenotypes. Linking them to genetic variability is essential for shorter and more efficient wheat breeding programs. A growing number of plant molecular information networks provide interlinked interoperable data to support the discovery of gene-phenotype interactions. A large body of scientific literature and observational data obtained in-field and under controlled conditions document wheat breeding experiments. The cross-referencing of this complementary information is essential. Text from databases and scientific publications has been identified early on as a relevant source of information. However, the wide variety of terms used to refer to traits and phenotype values makes it difficult to find and cross-reference the textual information, e.g. simple dictionary lookup methods miss relevant terms. Corpora with manually annotated examples are thus needed to evaluate and train textual information extraction methods. While several corpora contain annotations of human and animal phenotypes, no corpus is available for plant traits. This hinders the evaluation of text mining-based crop knowledge graphs (e.g. AgroLD, KnetMiner, WheatIS-FAIDARE) and limits the ability to train machine learning methods and improve the quality of information. The Triticum aestivum trait Corpus is a new gold standard for traits and phenotypes of wheat. It consists of 528 PubMed references that are fully annotated by trait, phenotype, and species. We address the interoperability challenge of crossing sparse assay data and publications by using the Wheat Trait and Phenotype Ontology to normalize trait mentions and the species taxonomy of the National Center for Biotechnology Information to normalize species. The paper describes the construction of the corpus. A study of the performance of state-of-the-art language models for both named entity recognition and linking tasks trained on the corpus shows that it is suitable for training and evaluation. This corpus is currently the most comprehensive manually annotated corpus for natural language processing studies on crop phenotype information from the literature.

Why it matches plant phenotyping methods小麦の形質・表現型抽出とエンティティ linking のための手動アノテーションコーパスを構築し、言語モデルの訓練・評価に用いており、表現型情報の取得手法とデータセットが研究の中心である。

titleA manually annotated text dataset for trait and phenotype extraction and entity linking in wheat breeding literature
Reproduction assets foundThe paper's core assets are publicly available: the TaeC annotated corpus (trait/phenotype/species annotations of 528 PubMed wheat references) on Recherche Data Gouv, the Wheat Trait and Phenotype Ontology on AgroPortal, the AlvisNLP bread wheat workflow on Forgemia, and the ToMap method code on GitHub, all with author
Dataset · publicThe corpus dataset TaeC is available under CC-BY-ND License at: https://entrepot.recherche.data.gouv.fr/dataset.xhtml?persistentId=doi:10.57745/GCYG3QOpen asset ↗entrepot.recherche.data.gouv.fr · doi:10.57745/GCYG3Qlines:142-152
Code · publicThe code of the ToMap method is available under Apache License at https://github.com/Bibliome/alvisnlp/tree/master/alvisnlp-bibliome/src/main/java/fr/inra/maiage/bibliome/alvisnlp/bibliomefactory/modules/tomapOpen asset ↗github.comlines:142-152
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published11 Jun 2024The Plant journal : for cell and molecular biologyCited by 9 · OpenAlex ↗

Genome-wide association study of stem structural characteristics that extracted by a high-throughput phenotypic analysis "LabelmeP rice" in rice.

RiceStem / branchMorphology / geometry measurementArchitecture / morphology / geometry

Stem is important for assimilating transport and plant strength; however, less is known about the genetic basis of its structural characteristics. In this study, a high-throughput method, "LabelmeP rice" was developed to generate 14 traits related to stem regions and vascular bundles, which allows the establishment of a stem cross-section phenotype dataset containing anatomical information of 1738 images from hand-cut transections of stems collected from 387 rice germplasm accessions grown over two successive seasons. Then, the phenotypic diversity of the rice accessions was evaluated. Genome-wide association studies identified 94, 83, and 66 significant single nucleotide polymorphisms (SNPs) for the assayed traits in 2 years and their best linear unbiased estimates, respectively. These SNPs can be integrated into 29 quantitative trait loci (QTL), and 11 of them were common in 2 years, while correlated traits shared 19. In addition, 173 candidate genes were identified, and six located at significant SNPs were repeatedly detected and annotated with a potential function in stem development. By using three introgression lines (chromosome segment substitution lines), four of the 29 QTLs were validated. LOC_Os01g70200, located on the QTL uq1.4, is detected for the area of small vascular bundles (SVB) and the rate of large vascular bundles number to SVB number. Besides, the CRISPR/Cas9 editing approach has elucidated the function of the candidate gene LOC_Os06g46340 in stem development. In conclusion, the results present a time- and cost-effective method that provides convenience for extracting rice stem anatomical traits and the candidate genes/QTL, which would help improve rice.

Why it matches plant phenotyping methodsイネ茎断面画像から解剖学的形質を抽出する高スループット手法を開発し、形質データセット構築と検証に用いており、表現型取得法が研究の中心である。

abstracta high-throughput method, "LabelmeP rice" was developed to generate 14 traits related to stem regions and vascular bundles
Reproduction assets foundThe paper's phenotyping/analysis Python code (LabelmeP rice pipeline) is explicitly provided as authors' supporting information Data S1, publicly available with the online article. Other referenced URLs (SNP-Seek, Q-TARO, RiceRNA, rmbreeding) are external databases, not paper-specific assets.
Code · publicThe Python code was provided as Data S1.Open asset ↗pdf-raw-page:10 lines:1-91