Advanced crop monitoring inside greenhouses is becoming one of the primary objectives of research centers. High-performance sensors, such as LiDAR or stereo cameras, have traditionally been employed for this purpose, though these often have a high cost. This work proposes a Visual-SLAM system using a monocular camera, which is significantly more cost-effective and specifically tailored for agricultural applications, such as mapping tomato crops in a greenhouse. Tests were carried out on a real tomato bunch, located in the Agroconnect experimental greenhouse. A ROS 2 Humble node was developed to run on the robot in order to capture images of these crops, which were then stored for offline processing. To generate a 3D mapped model for the crop in the greenhouse, the GLOMAP mapper, based on Structure-From-Motion, was integrated with the Hierarchical Localization toolbox. This initial mapping is a foundation for future, more advanced algorithms to analyze growth patterns, and optimize agricultural management. The system leverages a hierarchical localization paradigm based on a coarse-to-fine strategy: it first performs global retrieval to generate location hypotheses, then combines local features within the identified candidate regions. The results show a correct identification of the tomato cluster, correctly characterising the tomato that is occluded and inaccessible by classical vision technologies. The reconstructed 3D model was further validated against manual ground-truth measurements of fruit size, centroid position, and orientation, confirming the geometric accuracy of the proposed low-cost monocular pipeline.
Why it matches plant phenotyping methods単なる収穫対象の位置検出ではなく、単眼Visual-SLAMと3D再構成を開発し、果実サイズ・重心位置・向きを実測値で検証しているため、植物器官形質の取得手法が中心である。
abstractThis work proposes a Visual-SLAM system using a monocular camera, which is significantly more cost-effective and specifically tailored for agricultural applications, such as mapping tomato crops in a greenhouse.
Autonomous, cross-facility science requires capabilities that no individual project should have to build for itself: managed execution for long-lived services, versioned distribution of models to remote compute systems, governed access to large language models, a shared substrate for experimental data, and end-to-end provenance. The U.S. Department of Energy Genesis Mission platform, delivered through the American Science Cloud, provides these as reusable services. This paper reports how the Genesis platform enables cross-facility experiments and accelerates scientific discovery. We explore the plant phenotyping workflow of the Orchestrated Platform for Autonomous Laboratories as the exemplar: it couples Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory with the Frontier supercomputer. In a 40-day nickel-treatment campaign, the resulting workflow replaced roughly twelve hours of manual analysis with interactive queries returning in seconds to minutes.
Why it matches plant phenotyping methods植物フェノタイピングのワークフローを実証例として、実験施設とスーパーコンピュータを連携する再利用可能なプラットフォーム基盤を報告しており、フェノタイピング基盤が中心的である。
abstractWe explore the plant phenotyping workflow of the Orchestrated Platform for Autonomous Laboratories as the exemplar: it couples Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory with the Frontier supercomputer.
Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. We introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys. It contains 88,830 RGB images at $5280 \times 3956$ pixels, with a ground sampling distance of 3.6-5.8 mm, from 91 scenes spanning corn, soybean, wheat, and oat. Track A evaluates seven scene-optimized methods -- Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) variants -- on held-out views, photogrammetry-referenced depth, and canopy-height recovery. Track B tests four pretrained feed-forward models on zero-shot camera-pose and geometry estimation. The scene-optimized methods rank differently across the three targets: Splatfacto-big leads appearance, whereas Scaffold-GS leads depth and is statistically tied with Splatfacto for canopy height. Among feed-forward models, MapAnything leads on seven of the eight metrics, while the remaining models vary more across crops and fail severely on absolute scale in a way that alignment conceals. Repeated acquisitions reveal further sensitivities that differ by output type and by model, associated with position within the acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/
Why it matches plant phenotyping methods植物キャノピー高さという明示的な形質を対象に、UAV 3D再構成手法をベンチマークし、公開データセットとして提供しているため、フェノタイピング手法が中心である。
abstractWe introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys.
Reproduction assets foundThe paper introduces UAV3DCrop, a public benchmark of repeated multi-angle UAV crop surveys (88,830 RGB images, 91 scenes, four crops) with refined poses, photogrammetric depth references, and linked canopy-height and effective-LAI field measurements. The dataset is explicitly stated to be publicly available under CC BDataset · publiche acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/ .
Keywords:
UAV imagery; agricultural datasets; crop-field reconstruction; neural radiance fields;
Gaussian splatting; feed-forward geometry.
1 IntroductionOpen asset ↗UAV3DCroplines:1-90Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 15 Sept 2026
Tree growth determines how much CO2 is sequestered from the atmosphere and temporarily stored in woody biomass. At the same time tree growth is affected by increasing temperatures, more frequent drought periods, late frosts and other extreme events associated with climate change. While continuous measurements of radial (secondary) tree growth using dendrometers are well established, monitoring of shoot elongation (primary growth) has largely been neglected because suitable measurement techniques are lacking. As a result, the effects of climate change on primary tree growth remain insufficiently understood. This work aims at reconstructing native deciduous trees in 3D as a basis for measuring and monitoring shoot elongation over entire tree canopies. Here we explored the use of low-cost UAV photogrammetry and of a multi-camera CraneCam system under real-world conditions. Data were collected in two study areas over an entire growing season. We present sensor evaluations, photogrammetric data acquisition and processing strategies. A special focus is placed on the analysis of the resulting photogrammetric 3D point clouds in terms of accuracy, resolution and completeness. Results demonstrate 3D point accuracies of 5-6 mm for entire trees using consumer-grade UAVs weighing less than 250 g and a 3D reconstruction completeness between 92% and 98% depending on the UAV type. The paper introduces a novel 3Dprinted ground-truth branch to evaluate the capability to reconstructing fine-detail structures such as thin tree shoots. Finally, we discuss operational challenges and initial experiments towards a skeletonization of entire trees based on photogrammetric point clouds.
Why it matches plant phenotyping methods樹冠全体のシュート伸長を測定するための3D再構成手法を開発・評価し、センサー評価、取得・処理戦略、精度・完全性の検証を中心に扱っているため。
abstractThis work aims at reconstructing native deciduous trees in 3D as a basis for measuring and monitoring shoot elongation over entire tree canopies.
NeRF / 3D Gaussian SplattingLiDAR / point cloudLeafWhole plant / canopy / plot / field2D/3D reconstructionGrowth / time-series analysisTrackingArchitecture / morphology / geometryGrowth / development / phenology
Quantifying plant growth dynamics from sparse longitudinal 3D observations is fundamental for agriculture and plant sciences. Yet, plants pose unique challenges: they undergo intricate non-rigid deformations, exhibit changing topology as new organs emerge, and often lack explicit temporal correspondences between consecutive data acquisitions due to newly formed tissue. Methods designed for general scenes struggle to model topology changes and asynchronous organ growth characteristic of plants. To address these challenges, we introduce GrowFields, a compositional dynamic neural field representation for organ-aware 4D plant growth modelling from point cloud time series. Our approach decomposes a plant into its constituent organs and aligns each organ into its own canonical coordinate frame, isolating intrinsic growth patterns from global plant motion. We then learn a shared continuous neural deformation field that models temporal dynamics across all organs, conditioned on learnable per-organ latent codes capturing organ identity and growth characteristics. The resulting modular yet unified representation naturally accommodates the asynchronous development of plant organs while remaining grounded in the practical setting of organ-level plant tracking. We evaluate GrowFields on growth sequences from four plant species, assessing geometric fitting and organ tracking accuracy using manually annotated leaf-tip trajectories. Results demonstrate consistent improvements in spatial precision, temporal coherence, and morphological fidelity over a range of existing representations.
Why it matches plant phenotyping methods植物の3D時系列観測から器官追跡と成長動態を抽出する、オルガン認識型4Dニューラルフィールド手法の開発・評価が中心である。
abstractwe introduce GrowFields, a compositional dynamic neural field representation for organ-aware 4D plant growth modelling from point cloud time series.
Tomographic microscopy enables three-dimensional internal imaging but often requires expensive optical or X-ray instrumentation. Here we present an ultra-low-cost continuous-wave diffusive tomography (CWDT) system for biological samples. The system uses a smartphone microscope, a white LED coupled into an optical fiber, 3D-printed micropositioners, and a physics-based forward model optimized with machine learning. We demonstrate full-color volumetric reconstructions from a tartrazine-cleared poplar section, a scattering phantom, fungal mycelium near an Arabidopsis root, and thick poplar branch imaging with an inserted side-emitting fiber. The current results are qualitative and exploratory, but they show that scanned fiber illumination and inexpensive hardware can produce useful three-dimensional reconstruction outputs for low-cost microscopy experiments.
Why it matches plant phenotyping methods低コスト三次元断層イメージング法そのものを開発し、ポプラ組織・枝やシロイヌナズナ根近傍を対象に植物の内部構造を可視化しているため、植物形態の取得法として中心的です。
abstractHere we present an ultra-low-cost continuous-wave diffusive tomography (CWDT) system for biological samples.
Reproduction assets foundThe paper's raw imaging inputs, configurations, and reconstruction outputs for Figures 2–5 are publicly deposited on Kaggle. The analysis code repository is only 'prepared for release' (no confirmed public deposit yet), so it is listed as request-only. Hardware CAD mirrors are public but are instrument designs, not theDataset · publicFigure-level raw inputs, model configurations, selected outputs, and manifests are available through the Kaggle dataset https://www.kaggle.com/datasets/alingold/continuous-wave-diffusive-tomography .Open asset ↗continuous-wave-diffusive-tomographylines:108-129Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 11 Sept 2026
Commercial greenhouse cucumber production is graded by fruit length, which drives harvest scheduling, labour allocation, and logistics. Manual measurement with thread or caliper is accurate but infeasible at commercial scale. This paper presents CucumberVision, a non-contact length estimation framework using an Intel RealSense D435 RGB-D camera. A YOLO26n instance segmentation model locates cucumbers, and SAM (ViT-B backbone) refines each detection to a pixel-precise mask. Five methods are evaluated under matched conditions: (M1) a dominant-axis skeleton scan-line baseline; (M2) PCA on the bounding-box depth point cloud; (M3) SAM mask with medial-axis skeletonisation; (M4) a hybrid keypoint-guided approach using a YOLO26-pose model predicting five anatomical landmarks (KP0--KP4) with piecewise 3D arc-length; and (M5) a novel medial arc spline method fitting a cubic spline through the 3D medial axis of the SAM mask and computing arc length by trapezoidal integration -- the first such application to elongated vegetable measurement. All methods share five-frame burst depth averaging, colour-stream intrinsic alignment, and adaptive method selection with cascading fallbacks ensuring 100% coverage. A benchmark of 48 captures across seven cucumbers in three size categories (small ~8 cm, medium ~13 cm, large ~25 cm) with thread-based ground truth establishes a significant accuracy hierarchy: M1 (MAPE 9.68%) > M2 (5.31%) > M4 (5.51%) > M3 (5.82%) > M5 (4.13%). M5 significantly outperforms all competitors at Bonferroni-corrected alpha=0.0125. A secondary contribution is identifying a 12--18% length underestimation caused by using depth-stream rather than colour-stream intrinsics after rs.align(rs.stream.color) -- an under-reported error source. The complete system is released open source and runs in real time on a single consumer-grade GPU.
Why it matches plant phenotyping methodsRGB-D画像からキュウリ果実長を推定する手法を開発し、複数手法との比較検証、実測値によるベンチマーク、誤差要因分析まで行っており、植物形質取得が研究の中心です。
abstractThis paper presents CucumberVision, a non-contact length estimation framework using an Intel RealSense D435 RGB-D camera.
Alternatives to soil-based horticulture, such as hydroponics, have been developed to respond to food distribution concerns for dense urban centers. A new system was developed to track an individual lettuce plant's growth in a hydroponic environment, utilizing streams of measured information and available models to continuously update the growth trajectory estimates for a plant. These "digital twin" models were integrated into an operating hydroponic greenhouse, with custom horticultural and sensor hardware to grow and measure relevant information. To aid in updating model parameters, plant yield was continuously measured with a custom neural network, using RGB-D images of the plants as an input. The network, trained on a collected dataset of 1300 images, was able to estimate mass within 1.5 g of the ground-truth value. After integration into the custom system, digital twin growth projections could approximate future yield between one and four days in the future, maintaining around a 2 g forecasting error.
Why it matches plant phenotyping methodsRGB-D画像から個体レタスの収量・質量を推定するニューラルネットワークと、センサー統合型の成長追跡基盤が研究の中心であり、植物形質の取得・予測手法を実質的に開発・検証している。
abstractA new system was developed to track an individual lettuce plant's growth in a hydroponic environment, utilizing streams of measured information and available models to continuously update the growth trajectory estimates for a plant.
Accurate quantification of forest coverage and combustible biomass (fuel load) is critical for wildfire risk assessment and ecosystem management. However, traditional methods relying on airborne LiDAR or field surveys are cost-prohibitive and time-intensive, while satellite imagery often lacks the vertical resolution required for canopy volume analysis. This paper proposes a novel, automated pipeline for rapid forest inventory using virtual remote sensing data derived from Google Earth Studio (GES). Our approach first generates low-altitude orbital imagery and camera poses for a target region. For dense 3D reconstruction, we employ Pi-Long, developed within the VGGT-Long framework. This model serves as a scalable extension of the Pi-3 feed-forward Transformer architecture. To address the inherent scale ambiguity in monocular reconstruction, we introduce a metric recovery module that aligns the reconstructed trajectory with GES ground truth poses via Sim(3) Umeyama optimization. The metric-scale point cloud is then orthogonally projected into Bird's-Eye-View (BEV) height and density maps. Finally, we employ a watershed-based segmentation algorithm combined with height variance analysis to classify tree species (conifer vs. broadleaf), calculate Leaf Area Index (LAI), and estimate total fuel load. Experimental results demonstrate that this pipeline offers a scalable, cost-effective alternative to physical scanning, enabling near-real-time estimation of forest biomass with high geometric consistency.
Why it matches plant phenotyping methods森林の3D再構成、BEVマップ、分割・高さ分散解析を組み合わせ、LAIと燃料量という植物・群落形質を推定するパイプライン自体が中心的な技術貢献であるため。
abstractThis paper proposes a novel, automated pipeline for rapid forest inventory using virtual remote sensing data derived from Google Earth Studio (GES).
Natural and anthropogenic disturbances are impacting the health of forests worldwide. Monitoring forest disturbances at scale is important to inform conservation efforts. Here, we present a scalable approach for country-wide mapping of forest greenness anomalies at the 10 m resolution of Sentinel-2. Using relevant ecological and topographical context and an established representation of the vegetation cycle, we learn a predictive quantile model of the normalised difference vegetation index (NDVI) derived from Sentinel-2 data. The resulting expected seasonal cycles are used to detect NDVI anomalies across Switzerland between April 2017 and August 2025. Goodness-of-fit evaluations show that the conditional model explains 65% of the observed variations in the median seasonal cycle. The model consistently benefits from the local context information, particularly during the green-up period. The approach produces coherent spatial anomaly patterns and enables country-wide quantification of forest browning. Case studies with independent reference data from known events illustrate that the model reliably detects different types of disturbances.
Why it matches plant phenotyping methodsSentinel-2 NDVIを用いて森林キャノピーの季節変動から褐変・攪乱状態を推定する手法を開発し、適合度と独立参照データで検証しているため、単なる森林地図作成ではなく植物状態の取得・評価が中心である。
abstractwe present a scalable approach for country-wide mapping of forest greenness anomalies at the 10 m resolution of Sentinel-2.
Reproduction assets foundThe paper explicitly states that its code and interactive content are publicly available in the authors' GitHub repository. Other URLs in the article are cited third-party data sources (swisstopo, EnviDat, GDAL, TauDEM, WhiteboxTools) rather than paper-specific assets.Code · publicThe code and interactive content are available at https://github.com/SamanthaBiegel/s2-forest-browning-monitoring .Open asset ↗SamanthaBiegel/s2-forest-browning-monitoringlines:51-55Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 13 Sept 2026
Dense ground-truth disparity maps are practically unobtainable in forestry environments, where thin overlapping branches and complex canopy geometry defeat conventional depth sensors -- a critical bottleneck for training supervised stereo matching networks for autonomous UAV-based pruning. We present UE5-Forest, a photorealistic synthetic stereo dataset built entirely in Unreal Engine 5 (UE5). One hundred and fifteen photogrammetry-scanned trees from the Quixel Megascans library are placed in virtual scenes and captured by a simulated stereo rig whose intrinsics -- 63 mm baseline, 2.8 mm focal length, 3.84 mm sensor width -- replicate the ZED Mini camera mounted on our drone. Orbiting each tree at up to 2 m across three elevation bands (horizontal, +45 degrees, -45 degrees) yields 5,520 rectified 1920 x 1080 stereo pairs with pixel-perfect disparity labels. We provide a statistical characterisation of the dataset -- covering disparity distributions, scene diversity, and visual fidelity -- and a qualitative comparison with real-world Canterbury Tree Branches imagery that confirms the photorealistic quality and geometric plausibility of the rendered data. The dataset will be publicly released to provide the community with a ready-to-use benchmark and training resource for stereo-based forestry depth estimation.
Why it matches plant phenotyping methods樹木の枝・樹冠形状を対象とするステレオ深度推定データセットを開発し、画素単位の視差ラベルと実画像との比較検証を提供しており、植物構造の取得方法が中心である。
abstractWe present UE5-Forest, a photorealistic synthetic stereo dataset built entirely in Unreal Engine 5 (UE5).
Field / plotNeRF / 3D Gaussian SplattingPhotogrammetry / SfM / MVSLiDAR / point cloudLeafStem / branchWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionSkeletonization / topology
Saplings are key indicators of forest regeneration and overall forest health. However, their fine-scale architectural traits are difficult to capture with existing 3D sensing methods, which make quantitative evaluation difficult. Terrestrial Laser Scanners (TLS), Mobile Laser Scanners (MLS), or traditional photogrammetry approaches poorly reconstruct thin branches, dense foliage, and lack the scale consistency needed for long-term monitoring. Implicit 3D reconstruction methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) are promising alternatives, but cannot recover the true scale of a scene and lack any means to be accurately geo-localised. In this paper, we present a pipeline which fuses NeRF, LiDAR SLAM, and GNSS to enable repeatable, geo-localised ecological monitoring of saplings. Our system proposes a three-level representation: (i) coarse Earth-frame localisation using GNSS, (ii) LiDAR-based SLAM for centimetre-accurate localisation and reconstruction, and (iii) NeRF-derived object-centric dense reconstruction of individual saplings. This approach enables repeatable quantitative evaluation and long-term monitoring of sapling traits. Our experiments in forest plots in Wytham Woods (Oxford, UK) and Evo (Finland) show that stem height, branching patterns, and leaf-to-wood ratios can be captured with increased accuracy as compared to TLS. We demonstrate that accurate stem skeletons and leaf distributions can be measured for saplings with heights between 0.5m and 2m in situ, giving ecologists access to richer structural and quantitative data for analysing forest dynamics.
Why it matches plant phenotyping methodsNeRF・LiDAR SLAM・GNSSを融合した幼木の3D再構成・定位パイプラインを開発し、樹高、分枝、葉対木質比などの植物形質をTLSと比較検証しており、表現型取得手法が中心である。
abstractIn this paper, we present a pipeline which fuses NeRF, LiDAR SLAM, and GNSS to enable repeatable, geo-localised ecological monitoring of saplings.
Accurate per-branch 3D reconstruction is a prerequisite for autonomous UAV-based tree pruning; however, dense disparity maps from modern stereo matchers often remain too noisy for individual branch analysis in complex forest canopies. This paper introduces a progressive pipeline integrating DEFOM-Stereo foundation-model disparity estimation, SAM3 instance segmentation, and multi-stage depth optimization to deliver robust per-branch point clouds. Starting from a naive baseline, we systematically identify and resolve three error families through successive refinements. Mask boundary contamination is first addressed through morphological erosion and subsequently refined via a skeleton-preserving variant to safeguard thin-branch topology. Segmentation inaccuracy is then mitigated using LAB-space Mahalanobis color validation coupled with cross-branch overlap arbitration. Finally, depth noise - the most persistent error source - is initially reduced by outlier removal and median filtering, before being superseded by a robust five-stage scheme comprising MAD global detection, spatial density consensus, local MAD filtering, RGB-guided filtering, and adaptive bilateral filtering. Evaluated on 1920x1080 stereo imagery of Radiata pine (Pinus radiata) acquired with a ZED Mini camera (63 mm baseline) from a UAV in Canterbury, New Zealand, the proposed pipeline reduces the average per-branch depth standard deviation by 82% while retaining edge fidelity. The result is geometrically coherent 3D point clouds suitable for autonomous pruning tool positioning. All code and processed data are publicly released to facilitate further UAV forestry research.
Why it matches plant phenotyping methods樹木の枝を対象に、ステレオ深度推定・セグメンテーション・深度最適化による枝単位の3D形状抽出手法を開発・評価しており、植物器官の表現型取得が中心である。
abstractThis paper introduces a progressive pipeline integrating DEFOM-Stereo foundation-model disparity estimation, SAM3 instance segmentation, and multi-stage depth optimization to deliver robust per-branch point clouds.
Modeling the time-varying 3D appearance of plants during growth poses unique challenges: unlike most dynamic scenes, plants continuously generate new geometry as they expand, branch, and differentiate. Existing dynamic scene representations are ill-suited to this setting: deformation fields provide insufficient constraints to yield physically plausible scene dynamics, and 4D Gaussian splatting represents the same physical structures with different Gaussian primitives at different times, breaking temporal consistency. We introduce GrowFlow, a dynamic representation that couples 3D Gaussian primitives with a neural ordinary differential equation to model plant growth as a continuous flow field over geometric parameters (position, scale, and orientation). Our representation enables consistent appearance rendering and models nonlinear, continuous-time growth dynamics with full temporal correspondences for every primitive. To initialize a sufficient set of Gaussian primitives, we first reconstruct the mature plant and then learn a reverse-growth process, effectively simulating the plant's developmental history in reverse. GrowFlow achieves superior image quality and geometric coherence compared to prior methods on a new, multi-view timelapse dataset of plant growth, and provides the first temporally coherent representation for appearance modeling of growing 3D structures.
Why it matches plant phenotyping methods植物の成長を対象に、4D再構成と連続的な成長表現を開発し、幾何学的整合性と画像品質を既存手法・新規データセットで比較評価しているため、植物フェノタイピング手法が中心です。
abstractWe introduce GrowFlow, a dynamic representation that couples 3D Gaussian primitives with a neural ordinary differential equation to model plant growth as a continuous flow field over geometric parameters (position, scale, and orientation).
Aerial remote sensing efficiently surveys large areas, but accurate direct object-level measurement remains difficult in complex natural scenes. Advancements in 3D computer vision, particularly radiance field representations such as NeRF and 3D Gaussian splatting, can improve reconstruction fidelity from posed imagery. Nevertheless, direct aerial measurement of important attributes like tree diameter at breast height (DBH) remains challenging. Trunks in aerial forest scans are distant and sparsely observed in image views; at typical operating altitudes, stems may span only a few pixels. With these constraints, conventional reconstruction methods have inaccurate breast-height trunk geometry. TreeDGS is an aerial image reconstruction method that uses 3D Gaussian splatting as a continuous scene representation for trunk measurement. After SfM--MVS initialization and Gaussian optimization, we extract a dense point set from the Gaussian field using RaDe-GS's depth-aware cumulative-opacity integration and associate each sample with a multi-view opacity reliability score. Then, we isolate trunk points and estimate DBH using opacity-weighted solid-circle fitting. Evaluated on 10 plots with field-measured DBH, TreeDGS reaches 4.79 cm RMSE (about 2.6 pixels at this GSD) and outperforms a LiDAR baseline (7.66 cm RMSE). This shows that TreeDGS can enable accurate, low-cost aerial DBH measurement .
Why it matches plant phenotyping methods樹木のDBHという明示的な植物形質を、航空画像と3D Gaussian splattingから推定する新規手法を開発し、実測値およびLiDARと比較検証しているため、方法中心の植物フェノタイピング研究として含める。
abstractTreeDGS is an aerial image reconstruction method that uses 3D Gaussian splatting as a continuous scene representation for trunk measurement.
Semantic reconstruction of agricultural scenes plays a vital role in tasks such as phenotyping and yield estimation. However, traditional approaches based on manual scanning or fixed camera setups remain a major bottleneck, while active-mapping methods based solely on occupancy grids are too coarse for accurate trait estimation. To address this gap, we propose an active 3D reconstruction framework for horticultural environments using a mobile manipulator. The system integrates OctoMap with 3D Gaussian Splatting to enable accurate and efficient target-aware mapping. A low-resolution OctoMap provides probabilistic occupancy information for informative viewpoint selection and collision-free planning, while 3D Gaussian Splatting leverages geometric, photometric, and semantic information to optimize 3D Gaussians for high-fidelity scene reconstruction. We further introduce a robust mapping strategy that mitigates semantic segmentation and depth noise, together with a background pruning method that reduces memory and computational cost. We validate our framework across simulated, laboratory, and real greenhouse scenes, showing consistent improvements across three state-of-the-art Gaussian Splatting backbones. In simulation, where ground-truth geometry is available, our approach outperforms occupancy-based mapping in both reconstruction accuracy and runtime efficiency: compared with a 0.01m-resolution OctoMap, it doubles the fruit-level F1 score under noisy conditions while achieving up to a threefold reduction in runtime. Beyond simulation, novel-view synthesis quality also improves consistently in laboratory and real greenhouse environments, with PSNR and mIoU improving by up to 1.5 dB and 18%, respectively. Finally, the reconstructed semantic maps enable fruit counting and volume estimation with accuracies approaching 80%.
Why it matches plant phenotyping methods園芸ロボット向けの3D再構成・能動マッピング手法を開発し、果実の計数・体積推定という植物形質の取得に適用・検証しているため、フェノタイピング手法が中心的です。
titleOctoSplat: Hybrid OctoMap-Gaussian Splatting for Active Semantic Mapping and Phenotyping with Horticultural Robots
Reproduction assets foundThe paper's supplementary material is hosted on the authors' public project page (jrcuaranv.github.io/octosplat), and the authors state that all code and data are publicly available. The SimSense repository is a third-party depth-sensor simulator tool, not a paper-specific asset.Code · publicAll code and data are publicly available to facilitate reproducibility.Open asset ↗lines:59-163Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 13 Sept 2026
High-resolution UAV photogrammetry has become a key technology for precision agriculture, enabling centimeter-level crop monitoring and point-level plant localization. However, point-level maize localization in UAV imagery remains challenging due to (1) extremely small object-to-pixel ratios, typically less than 0.1%, (2) prohibitive computational costs of quadratic attention on ultra-high-resolution images larger than 3000 x 4000 pixels, and (3) agricultural scene-specific complexities such as sparse object distribution and environmental variability that are poorly handled by general-purpose vision models. To address these challenges, we propose the Additive Kolmogorov-Arnold Transformer (AKT), which replaces conventional multilayer perceptrons with Pade Kolmogorov-Arnold Network (PKAN) modules to enhance functional expressivity for small-object feature extraction, and introduces PKAN Additive Attention (PAA) to model multiscale spatial dependencies with reduced computational complexity. In addition, we present the Point-based Maize Localization (PML) dataset, consisting of 1,928 high-resolution UAV images with approximately 501,000 point annotations collected under real field conditions. Extensive experiments show that AKT achieves an average F1-score of 62.8%, outperforming state-of-the-art methods by 4.2%, while reducing FLOPs by 12.6% and improving inference throughput by 20.7%. For downstream tasks, AKT attains a mean absolute error of 7.1 in stand counting and a root mean square error of 1.95-1.97 cm in interplant spacing estimation. These results demonstrate that integrating Kolmogorov-Arnold representation theory with efficient attention mechanisms offers an effective framework for high-resolution agricultural remote sensing.
Why it matches plant phenotyping methodsUAV画像から個体位置を抽出する手法を開発し、個体数と株間距離という植物群落形質を推定しており、データセット構築と技術評価も中心的である。
abstractTo address these challenges, we propose the Additive Kolmogorov-Arnold Transformer (AKT)
AppleCottonPearField / plotNeRF / 3D Gaussian SplattingFruitCounting2D/3D reconstructionSegmentation
Rigorous crop counting is crucial for effective agricultural management and informed intervention strategies. However, in outdoor field environments, partial occlusions combined with inherent ambiguity in distinguishing clustered crops from individual viewpoints poses an immense challenge for image-based segmentation methods. To address these problems, we introduce a novel crop counting framework designed for exact enumeration via 3D instance segmentation. Our approach utilizes 2D images captured from multiple viewpoints and associates independent instance masks for neural radiance field (NeRF) view synthesis. We introduce crop visibility and mask consistency scores, which are incorporated alongside 3D information from a NeRF model. This results in an effective segmentation of crop instances in 3D and highly-accurate crop counts. Furthermore, our method eliminates the dependence on crop-specific parameter tuning. We validate our framework on three agricultural datasets consisting of cotton bolls, apples, and pears, and demonstrate consistent counting performance despite major variations in crop color, shape, and size. A comparative analysis against the state of the art highlights superior performance on crop counting tasks. Lastly, we contribute a cotton plant dataset to advance further research on this topic.
Why it matches plant phenotyping methodsNeRFと3Dインスタンスセグメンテーションを用いて作物個体・器官数を推定する画像ベース表現型計測手法を開発・検証しており、方法が研究の中心である。
abstractwe introduce a novel crop counting framework designed for exact enumeration via 3D instance segmentation.
Reproduction assets foundThe paper contributes a public infield cotton plant dataset (8 plants, ~150 iPhone images each, ground-truth boll counts, SAM instance masks) and states that source code, dataset, and multimedia are available at the authors' public project page, which is an allowed URL. The spectacularai GitHub URL is a generic third-pDataset · publicthat incorporates crop visibility
and mask consistency, enabling robustness against occlusions and annotation
discrepancies.
•
We release a public infield cotton plant dataset designed for 3D
rendering and cotton boll counting tasks.
The source code, dataset, and multimedia material associated with this project
can be found at
https://robotic-vision-lab.github.io/cropnerf .
II Related Work
II-A Image-Based Techniques
Image-based methods typically employ object detection to identify crops within
images. For example, Chen et al. [ 4 ] utilized multiple
convolutional neural networks (CNNs) to map input images to total fruit counts.
Similarly, Häni et al. [ 5 ] formulated crop counting as a
multOpen asset ↗lines:108-187Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Potato yield is a key indicator for optimizing cultivation practices in agriculture. Potato yield can be estimated on harvesters using RGB-D cameras, which capture three-dimensional (3D) information of individual tubers moving along the conveyor belt. However, point clouds reconstructed from RGB-D images are incomplete due to self-occlusion, leading to systematic underestimation of tuber weight. To address this, we introduce PointRAFT, a high-throughput point cloud regression network that directly predicts continuous 3D shape properties, such as tuber weight, from partial point clouds. Rather than reconstructing full 3D geometry, PointRAFT infers target values directly from raw 3D data. Its key architectural novelty is an object height embedding that incorporates tuber height as an additional geometric cue, improving weight prediction under practical harvesting conditions. PointRAFT was trained and evaluated on 26,688 partial point clouds collected from 859 potato tubers across four cultivars and three growing seasons on an operational harvester in Japan. On a test set of 5,254 point clouds from 172 tubers, PointRAFT achieved a mean absolute error of 12.0 g and a root mean squared error of 17.2 g, substantially outperforming a linear regression baseline and a standard PointNet++ regression network. With an average inference time of 6.3 ms per point cloud, PointRAFT supports processing rates of up to 150 tubers per second, meeting the high-throughput requirements of commercial potato harvesters. Beyond potato weight estimation, PointRAFT provides a versatile regression network applicable to a wide range of 3D phenotyping and robotic perception tasks. The code, network weights, and a subset of the dataset are publicly available at https://github.com/pieterblok/pointraft.git.
Why it matches plant phenotyping methods部分点群からジャガイモ塊茎重量を推定する3D深層学習手法を開発・評価しており、植物形質取得が研究の中心である。
abstractwe introduce PointRAFT, a high-throughput point cloud regression network that directly predicts continuous 3D shape properties, such as tuber weight, from partial point clouds.
Reproduction assets foundThe paper publicly releases its authors' analysis code and trained network weights on GitHub, and a subset of its potato tuber partial point cloud dataset (with ground truth weights) on Hugging Face. Both are paper-specific, public, and actionable.Code · publicThe code, network weights, and a subset of the dataset are publicly available at https://github.com/pieterblok/pointraft.git .Open asset ↗pieterblok/pointraftlines:1-93Dataset · publicA subset of the datasets generated and/or analyzed during this study is publicly available at: https://huggingface.co/datasets/UTokyo-FieldPhenomics-Lab/3DPotatoTwinOpen asset ↗UTokyo-FieldPhenomics-Lab/3DPotatoTwinlines:447-463Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Automating tasks in orchards is challenging because of the large amount of variation in the environment and occlusions. One of the challenges is apple pose estimation, where key points, such as the calyx, are often occluded. Recently developed pose estimation methods no longer rely on these key points, but still require them for annotations, making annotating challenging and time-consuming. Due to the abovementioned occlusions, there can be conflicting and missing annotations of the same fruit between different images. Novel 3D reconstruction methods can be used to simplify annotating and enlarge datasets. We propose a novel pipeline consisting of 3D Gaussian Splatting to reconstruct an orchard scene, simplified annotations, automated projection of the annotations to images, and the training and evaluation of a pose estimation method. Using our pipeline, 105 manual annotations were required to obtain 28,191 training labels, a reduction of 99.6%. Experimental results indicated that training with labels of fruits that are $\leq95\%$ occluded resulted in the best performance, with a neutral F1 score of 0.927 on the original images and 0.970 on the rendered images. Adjusting the size of the training dataset had small effects on the model performance in terms of F1 score and pose estimation accuracy. It was found that the least occluded fruits had the best position estimation, which worsened as the fruits became more occluded. It was also found that the tested pose estimation method was unable to correctly learn the orientation estimation of apples.
Why it matches plant phenotyping methodsリンゴの姿勢推定アノテーションを大幅に効率化する3D再構成・自動ラベル投影パイプラインを開発し、姿勢推定性能も評価しているため、植物フェノタイピング手法が中心である。
abstractWe propose a novel pipeline consisting of 3D Gaussian Splatting to reconstruct an orchard scene, simplified annotations, automated projection of the annotations to images, and the training and evaluation of a pose estimation method.
Reproduction assets foundThe paper explicitly provides two paper-specific public assets: the authors' phenotyping/pose-estimation pipeline code on GitHub and the collected apple orchard image dataset on a 4TU DOI. Both are directly used for the paper's measurements and analysis.Dataset · publicIn total, 367 images were collected. The dataset is available at https://doi.org/10.4121/976c94f2-028f-4291-adfd-20eb82b0f647Open asset ↗10.4121/976c94f2-028f-4291-adfd-20eb82b0f647lines:92-108Plant phenotyping relevance match · UnverifiedarXiv · checked 6 Sept 2026
This extended abstract details our solution for the Global Wheat Full Semantic Segmentation Competition. We developed a systematic self-training framework. This framework combines a two-stage hybrid training strategy with extensive data augmentation. Our core model is SegFormer with a Mix Transformer (MiT-B4) backbone. We employ an iterative teacher-student loop. This loop progressively refines model accuracy. It also maximizes data utilization. Our method achieved competitive performance. This was evident on both the Development and Testing Phase datasets.
Why it matches plant phenotyping methods小麦穂を画像から分割する手法の開発が中心であり、植物器官の表現型取得に直接関係する。
titlePseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
Lychee is a high-value subtropical fruit. The adoption of vision-based harvesting robots can significantly improve productivity while reduce reliance on labor. High-quality data are essential for developing such harvesting robots. However, there are currently no consistently and comprehensively annotated open-source lychee datasets featuring fruits in natural growing environments. To address this, we constructed a dataset to facilitate lychee detection and maturity classification. Color (RGB) images were acquired under diverse weather conditions, and at different times of the day, across multiple lychee varieties, such as Nuomici, Feizixiao, Heiye, and Huaizhi. The dataset encompasses three different ripeness stages and contains 11,414 images, consisting of 878 raw RGB images, 8,780 augmented RGB images, and 1,756 depth images. The images are annotated with 9,658 pairs of lables for lychee detection and maturity classification. To improve annotation consistency, three individuals independently labeled the data, and their results were then aggregated and verified by a fourth reviewer. Detailed statistical analyses were done to examine the dataset. Finally, we performed experiments using three representative deep learning models to evaluate the dataset. It is publicly available for academic
Why it matches plant phenotyping methodsライチ果実の成熟段階という植物器官の状態をRGB-D画像から分類するデータセットを構築し、アノテーション検証と深層学習モデル評価を行っており、表現型取得・評価手法が中心である。
abstractwe constructed a dataset to facilitate lychee detection and maturity classification.
Reproduction assets foundThe authors publicly release the paper's lychee RGB-D image dataset (raw/augmented RGB images, depth maps, detection and maturity annotations) and the Python scripts for data augmentation, image similarity comparison, and annotation in the same GitHub repository.Dataset · publicchees, the non-augmented models
produced misclassifications with lower recognition and accuracy, whereas the augmented models
avoided these issues. Overall, the results demonstrate that the data augmentation method effectively
improves the comprehensive performance of the models.
5. Data Availability
The dataset is available at:https://github.com/SeiriosLab/Lychee. The Python scripts for data
augmentation, image similarity comparison, and annotation are available within the same
repository under the tree/main/script directory.Open asset ↗SeiriosLab/Lycheepdf-raw-page:13 lines:1-55Plant phenotyping relevance match · UnverifiedarXiv · checked 13 Sept 2026
Learning 3D parametric shape models of objects has gained popularity in vision and graphics and has showed broad utility in 3D reconstruction, generation, understanding, and simulation. While powerful models exist for humans and animals, equally expressive approaches for modeling plants are lacking. In this work, we present Demeter, a data-driven parametric model that encodes key factors of a plant morphology, including topology, shape, articulation, and deformation into a compact learned representation. Unlike previous parametric models, Demeter handles varying shape topology across various species and models three sources of shape variation: articulation, subcomponent shape variation, and non-rigid deformation. To advance crop plant modeling, we collected a large-scale, ground-truthed dataset from a soybean farm as a testbed. Experiments show that Demeter effectively synthesizes shapes, reconstructs structures, and simulates biophysical processes. Code and data is available at https://tianhang-cheng.github.io/Demeter/.
Why it matches plant phenotyping methods植物形態のトポロジー・形状・関節・変形を表現する3Dパラメトリックモデルを開発し、実世界の作物データセットで再構成・シミュレーションを評価しているため、形態フェノタイピング手法が中心である。
abstractwe present Demeter, a data-driven parametric model that encodes key factors of a plant morphology, including topology, shape, articulation, and deformation into a compact learned representation.
3D phenotyping of plants plays a crucial role for understanding plant growth, yield prediction, and disease control. We present a pipeline capable of generating high-quality 3D reconstructions of individual agricultural plants. To acquire data, a small commercially available UAV captures images of a selected plant. Apart from placing ArUco markers, the entire image acquisition process is fully autonomous, controlled by a self-developed Android application running on the drone's controller. The reconstruction task is particularly challenging due to environmental wind and downwash of the UAV. Our proposed pipeline supports the integration of arbitrary state-of-the-art 3D reconstruction methods. To mitigate errors caused by leaf motion during image capture, we use an iterative method that gradually adjusts the input images through deformation. Motion is estimated using optical flow between the original input images and intermediate 3D reconstructions rendered from the corresponding viewpoints. This alignment gradually reduces scene motion, resulting in a canonical representation. After a few iterations, our pipeline improves the reconstruction of state-of-the-art methods and enables the extraction of high-resolution 3D meshes. We will publicly release the source code of our reconstruction pipeline. Additionally, we provide a dataset consisting of multiple plants from various crops, captured across different points in time.
Why it matches plant phenotyping methodsUAV画像から植物個体の高解像度3D形状を再構成する手法と、風による葉の動きを補正する技術を開発しており、植物表現型取得が中心です。データセット提供も含みます。
abstractWe present a pipeline capable of generating high-quality 3D reconstructions of individual agricultural plants.
AI-driven crop health mapping systems offer substantial advantages over conventional monitoring approaches through accelerated data acquisition and cost reduction. However, widespread farmer adoption remains constrained by technical limitations in orthomosaic generation from sparse aerial imagery datasets. Traditional photogrammetric reconstruction requires 70-80\% inter-image overlap to establish sufficient feature correspondences for accurate geometric registration. AI-driven systems operating under resource-constrained conditions cannot consistently achieve these overlap thresholds, resulting in degraded reconstruction quality that undermines user confidence in autonomous monitoring technologies. In this paper, we present Ortho-Fuse, an optical flow-based framework that enables the generation of a reliable orthomosaic with reduced overlap requirements. Our approach employs intermediate flow estimation to synthesize transitional imagery between consecutive aerial frames, artificially augmenting feature correspondences for improved geometric reconstruction. Experimental validation demonstrates a 20\% reduction in minimum overlap requirements. We further analyze adoption barriers in precision agriculture to identify pathways for enhanced integration of AI-driven monitoring systems.
Why it matches plant phenotyping methods作物の健康状態マップ作成を目的とする航空画像のオルソモザイク生成法を開発し、重複率低減を実験検証しており、画像取得・再構成手法が中心である。
titleOrtho-Fuse: Orthomosaic Generation for Sparse High-Resolution Crop Health Maps Through Intermediate Optical Flow Estimation
Reproduction assets foundThe paper explicitly states that the authors' code and dataset (aerial imagery used for orthomosaic generation and crop health analysis) are publicly available at the project page https://rugvedkatole.github.io/OrthoFUSE/, which is an allowed URL. This qualifies as a paper-specific public asset containing the authors' Code · publicThe code and dataset are available at https://rugvedkatole.github.io/OrthoFUSE/Open asset ↗https://rugvedkatole.github.io/OrthoFUSE/lines:1-54Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 13 Sept 2026
Repeated plant monitoring is essential for tracking crop growth, and 3D reconstruction enables consistent comparison across monitoring sessions. However, rebuilding a 3D model from scratch in every session is costly and overlooks informative geometry already observed previously. We propose efficient view planning guided by a previous-session reconstruction, which reuses a 3D model from the previous session to improve active perception in the current session. Based on this previous-session reconstruction, our method replaces iterative next-best-view planning with one-shot view planning that selects an informative set of views and computes the globally shortest execution path connecting them. Experiments on real multi-session datasets, including public single-plant scans and a newly collected greenhouse crop-row dataset, show that our method achieves comparable or higher surface coverage with fewer executed views and shorter robot paths than iterative and one-shot baselines.
Why it matches plant phenotyping methods植物の反復モニタリング向けに、過去セッションの3D再構成を利用した視点計画とロボット撮影経路を開発・評価しており、植物形状取得の方法が中心です。
abstractWe propose efficient view planning guided by a previous-session reconstruction, which reuses a 3D model from the previous session to improve active perception in the current session.
Photogrammetry / SfM / MVSLiDAR / point cloudCalibration / preprocessingSegmentation
Accurate point cloud segmentation for plant organs is crucial for 3D plant phenotyping. Existing solutions are designed problem-specific with a focus on certain plant species or specified sensor-modalities for data acquisition. Furthermore, it is common to use extensive pre-processing and down-sample the plant point clouds to meet hardware or neural network input size requirements. We propose a simple, yet effective algorithm KDSS for sub-sampling of biological point clouds that is agnostic to sensor data and plant species. The main benefit of this approach is that we do not need to down-sample our input data and thus, enable segmentation of the full-resolution point cloud. Combining KD-SS with current state-of-the-art segmentation models shows satisfying results evaluated on different modalities such as photogrammetry, laser triangulation and LiDAR for various plant species. We propose KD-SS as lightweight resolution-retaining alternative to intensive pre-processing and down-sampling methods for plant organ segmentation regardless of used species and sensor modality.
Why it matches plant phenotyping methods植物器官の3D点群セグメンテーションと、センサー・種に依存しないサブサンプリング手法を開発・評価しており、植物フェノタイピングのための形態抽出手法が中心である。
abstractAccurate point cloud segmentation for plant organs is crucial for 3D plant phenotyping.
Street trees are vital to urban livability, providing ecological and social benefits. Establishing a detailed, accurate, and dynamically updated street tree inventory has become essential for optimizing these multifunctional assets within space-constrained urban environments. Given that traditional field surveys are time-consuming and labor-intensive, automated surveys utilizing Mobile Mapping Systems (MMS) offer a more efficient solution. However, existing MMS-acquired tree datasets are limited by small-scale scene, limited annotation, or single modality, restricting their utility for comprehensive analysis. To address these limitations, we introduce WHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset. Collected across two distinct cities, WHU-STree integrates synchronized point clouds and high-resolution images, encompassing 21,007 annotated tree instances across 50 species and 2 morphological parameters. Leveraging the unique characteristics, WHU-STree concurrently supports over 10 tasks related to street tree inventory. We benchmark representative baselines for two key tasks--tree species classification and individual tree segmentation. Extensive experiments and in-depth analysis demonstrate the significant potential of multi-modal data fusion and underscore cross-domain applicability as a critical prerequisite for practical algorithm deployment. In particular, we identify key challenges and outline potential future works for fully exploiting WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV/WHU-STree.
Why it matches plant phenotyping methods樹木の個体セグメンテーションと形態パラメータを含むマルチモーダルデータセットを構築し、ベンチマークする研究であり、植物個体の状態・形態抽出手法が中心である。
abstractWHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset.
Reproduction assets foundThe paper's core asset is the WHU-STree multi-modal street tree dataset (point clouds, panoramic images, 21,007 annotated tree instances, 50 species, height/DBH), which the authors state is publicly accessible via their GitHub organization WHU-USI3DV. The Zenodo DOIs in the reference list belong to cited prior datasetsDataset · publicticular, we
identify key challenges and outline potential future works for fully exploit-
ing WHU-STree, encompassing multi-modal fusion, multi-task collaboration,
cross-domain generalization, spatial pattern learning, and Multi-modal Large
Language Model for street tree asset management. The WHU-STree dataset
is accessible at: https://github.com/WHU-USI3DV /WHU-STree.
Keywords: Deep learning, Tree inventory, Individual tree segmentation,
Tree species classification, Multi-modal, Mobile mapping system
1. Introduction
Street trees, vital to urban ecosystems, provide ecological benefits (e.g.,
shade (Kumar et al., 2024), air purification (Grundstrém and Pleijel, 2014),
noise reductiOpen asset ↗WHU-STreepdf-raw-page:2 lines:1-35Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Cotton is one of the most important natural fiber crops worldwide, yet harvesting remains limited by labor-intensive manual picking, low efficiency, and yield losses from missing the optimal harvest window. Accurate recognition of cotton bolls and their maturity is therefore essential for automation, yield estimation, and breeding research. We propose Cott-ADNet, a lightweight real-time detector tailored to cotton boll and flower recognition under complex field conditions. Building on YOLOv11n, Cott-ADNet enhances spatial representation and robustness through improved convolutional designs, while introducing two new modules: a NeLU-enhanced Global Attention Mechanism to better capture weak and low-contrast features, and a Dilated Receptive Field SPPF to expand receptive fields for more effective multi-scale context modeling at low computational cost. We curate a labeled dataset of 4,966 images, and release an external validation set of 1,216 field images to support future research. Experiments show that Cott-ADNet achieves 91.5% Precision, 89.8% Recall, 93.3% mAP50, 71.3% mAP, and 90.6% F1-Score with only 7.5 GFLOPs, maintaining stable performance under multi-scale and rotational variations. These results demonstrate Cott-ADNet as an accurate and efficient solution for in-field deployment, and thus provide a reliable basis for automated cotton harvesting and high-throughput phenotypic analysis. Code and dataset is available at https://github.com/SweefongWong/Cott-ADNet.
Why it matches plant phenotyping methods綿花の花・ボール認識を対象とする画像解析手法を開発し、データセット作成、外部検証、性能評価まで行っており、植物器官の表現型取得が中心である。
abstractWe propose Cott-ADNet, a lightweight real-time detector tailored to cotton boll and flower recognition under complex field conditions.
Reproduction assets foundThe paper explicitly states that its code and curated cotton boll/flower detection dataset (4,966 labeled images plus a 1,216-image external validation set) are publicly released at the authors' GitHub repository. The ultralytics repository is a generic third-party library, not a paper-specific asset.Code · publicy 7.5 GFLOPs, maintaining stable performance under multi-scale and rotational variations. These results demonstrate Cott-ADNet as an accurate and efficient solution for in-field deployment, and thus provide a reliable basis for automated cotton harvesting and high-throughput phenotypic analysis. Code and dataset is available at https://github.com/SweefongWong/Cott-ADNet .
† † footnotetext: ∗ * Corresponding author: cuij@wfu.edu
Index Terms :
cotton, cotton boll detection, lightweight object detection, rotational convolution
1 Introduction
Cotton is one of the most critical economic crops worldwide, accounting for nearly 35% of global natural fiber production. It underpins industries such as Open asset ↗SweefongWong/Cott-ADNetlines:1-57Plant phenotyping relevance match · UnverifiedarXiv · checked 15 Sept 2026
Field / plotMultispectral / hyperspectralPhysiological trait estimationPhotosynthesis / fluorescence
Sun-induced fluorescence (SIF) as a close remote sensing based proxy for photosynthesis is accepted as a useful measure to remotely monitor vegetation health and gross primary productivity. In this work we present the new retrieval method WAFER (WAvelet decomposition FluorEscence Retrieval) based on wavelet decompositions of the measured spectra of reflected radiance as well as a reference radiance not containing fluorescence. By comparing absolute absorption line depths by means of the corresponding wavelet coefficients, a relative reflectance is retrieved independently of the fluorescence, i.e. without introducing a coupling between reflectance and fluorescence. The fluorescence can then be derived as the remaining offset. This method can be applied to arbitrary chosen wavelength windows in the whole spectral range, such that all the spectral data available is exploited, including the separation into several frequency (i.e. width of absorption lines) levels and without the need of extensive training datasets. At the same time, the assumptions about the reflectance shape are minimal and no spectral shape assumptions are imposed on the fluorescence, which not only avoids biases arising from wrong or differing fluorescence models across different spatial scales and retrieval methods but also allows for the exploration of this spectral shape for different measurement setups. WAFER is tested on a synthetic dataset as well as several diurnal datasets acquired with a field spectrometer (FloX) over an agricultural site. We compare the WAFER method to two established retrieval methods, namely the improved Fraunhofer line discrimination (iFLD) method and spectral fitting method (SFM) and find a good agreement with the added possibility of exploring the true spectral shape of the offset signal and free choice of the retrieval window. (abbreviated)
Why it matches plant phenotyping methods植物の光合成状態に関連するSIFを分光データから抽出する新手法を開発し、合成・実測データで既存手法と比較検証しており、表現型取得手法が中心である。
abstractIn this work we present the new retrieval method WAFER (WAvelet decomposition FluorEscence Retrieval)
StrawberryField / plotLiDAR / point cloudFlowerObject detectionPose / keypoint estimation
The small scale of urban farms and the commercial availability of low-cost robots (such as the FarmBot) that automate simple tending tasks enable an accessible platform for plant phenotyping. We have used a FarmBot with a custom camera end-effector to estimate strawberry plant flower pose (for robotic pollination) from acquired 3D point cloud models. We describe a novel algorithm that translates individual occupancy grids along orthogonal axes of a point cloud to obtain 2D images corresponding to the six viewpoints. For each image, 2D object detection models for flowers are used to identify 2D bounding boxes which can be converted into the 3D space to extract flower point clouds. Pose estimation is performed by fitting three shapes (superellipsoids, paraboloids and planes) to the flower point clouds and compared with manually labeled ground truth. Our method successfully finds approximately 80% of flowers scanned using our customized FarmBot platform and has a mean flower pose error of 7.7 degrees, which is sufficient for robotic pollination and rivals previous results. All code will be made available at https://github.com/harshmuriki/flowerPose.git.
Why it matches plant phenotyping methodsカスタムカメラ付きロボットによる3D花姿勢推定アルゴリズムとプラットフォームを開発・評価しており、花の姿勢という植物形質の取得が中心である。
abstractenable an accessible platform for plant phenotyping
Reproduction assets foundThe paper's flower pose estimation pipeline (translating occupancy grid, 2D/3D conversion, shape fitting) has an explicit authors' code deposit statement with a public GitHub URL, phrased as future availability ('will be made available'), so actionability is likely but not fully confirmed. No public dataset of the FarmCode · publiclower point clouds and compared with manually labeled ground truth. Our method successfully finds approximately 80% of flowers scanned using our customized FarmBot platform and has a mean flower pose error of 7.7 degrees, which is sufficient for robotic pollination and rivals previous results. All code will be made available at https://github.com/harshmuriki/flowerPose.git .
I Introduction
Urban farms [ 1 ] provide healthy food to local communities and can serve as platforms for education and sustainability. Unlike their rural counterparts, urban farms are usually small in scale and commercially available robotic systems such as the FarmBot [ 2 ] have been developed to help automate basic cuOpen asset ↗harshmuriki/flowerPoselines:1-53Plant phenotyping relevance match · UnverifiedarXiv · checked 13 Sept 2026
Digital twin applications offered transformative potential by enabling real-time monitoring and robotic simulation through accurate virtual replicas of physical assets. The key to these systems is 3D reconstruction with high geometrical fidelity. However, existing methods struggled under field conditions, especially with sparse and occluded views. This study developed a two-stage framework (DATR) for the reconstruction of apple trees from sparse views. The first stage leverages onboard sensors and foundation models to semi-automatically generate tree masks from complex field images. Tree masks are used to filter out background information in multi-modal data for the single-image-to-3D reconstruction at the second stage. This stage consists of a diffusion model and a large reconstruction model for respective multi view and implicit neural field generation. The training of the diffusion model and LRM was achieved by using realistic synthetic apple trees generated by a Real2Sim data generator. The framework was evaluated on both field and synthetic datasets. The field dataset includes six apple trees with field-measured ground truth, while the synthetic dataset featured structurally diverse trees. Evaluation results showed that our DATR framework outperformed existing 3D reconstruction methods across both datasets and achieved domain-trait estimation comparable to industrial-grade stationary laser scanners while improving the throughput by $\sim$360 times, demonstrating strong potential for scalable agricultural digital twin systems.
Why it matches plant phenotyping methodsリンゴ樹の疎視点画像から3D形状を再構成し、樹体形質を推定する手法の開発と、圃場・合成データでの評価が研究の中心であるため。
abstractThis study developed a two-stage framework (DATR) for the reconstruction of apple trees from sparse views.
RootObject detection2D/3D reconstructionSkeletonization / topologyRoot system architecture
Plant roots typically exhibit a highly complex and dense architecture, incorporating numerous slender lateral roots and branches, which significantly hinders the precise capture and modeling of the entire root system. Additionally, roots often lack sufficient texture and color information, making it difficult to identify and track root traits using visual methods. Previous research on roots has been largely confined to 2D studies; however, exploring the 3D architecture of roots is crucial in botany. Since roots grow in real 3D space, 3D phenotypic information is more critical for studying genetic traits and their impact on root development. We have introduced a 3D root skeleton extraction method that efficiently derives the 3D architecture of plant roots from a few images. This method includes the detection and matching of lateral roots, triangulation to extract the skeletal structure of lateral roots, and the integration of lateral and primary roots. We developed a highly complex root dataset and tested our method on it. The extracted 3D root skeletons showed considerable similarity to the ground truth, validating the effectiveness of the model. This method can play a significant role in automated breeding robots. Through precise 3D root structure analysis, breeding robots can better identify plant phenotypic traits, especially root structure and growth patterns, helping practitioners select seeds with superior root systems. This automated approach not only improves breeding efficiency but also reduces manual intervention, making the breeding process more intelligent and efficient, thus advancing modern agriculture.
Why it matches plant phenotyping methods植物根系の3D骨格・構造という表現型を画像から抽出する手法を開発し、データセット上で検証しており、方法が研究の中心である。
abstractWe have introduced a 3D root skeleton extraction method that efficiently derives the 3D architecture of plant roots from a few images.
StrawberryTomatoNeRF / 3D Gaussian SplattingFruit2D/3D reconstructionSegmentation
DexFruit is a robotic manipulation framework that enables gentle, autonomous handling of fragile fruit and precise evaluation of damage. Many fruits are fragile and prone to bruising, thus requiring humans to manually harvest them with care. In this work, we demonstrate by using optical tactile sensing, autonomous manipulation of fruit with minimal damage can be achieved. We show that our tactile informed diffusion policies outperform baselines in both reduced bruising and pick-and-place success rate across three fruits: strawberries, tomatoes, and blackberries. In addition, we introduce FruitSplat, a novel technique to represent and quantify visual damage in high-resolution 3D representation via 3D Gaussian Splatting (3DGS). Existing metrics for measuring damage lack quantitative rigor or require expensive equipment. With FruitSplat, we distill a 2D strawberry mask as well as a 2D bruise segmentation mask into the 3DGS representation. Furthermore, this representation is modular and general, compatible with any relevant 2D model. Overall, we demonstrate a 92% grasping policy success rate, up to a 20% reduction in visual bruising, and up to an 31% improvement in grasp success rate on challenging fruit compared to our baselines across our three tested fruits. We rigorously evaluate this result with over 630 trials. Please checkout our website at https://dex-fruit.github.io .
Why it matches plant phenotyping methodsFruitSplatは果実の外観損傷・打撲を3D表現として定量化する画像ベースの植物表現型手法であり、手法開発と大規模な技術評価が中心である。
abstractwe introduce FruitSplat, a novel technique to represent and quantify visual damage in high-resolution 3D representation via 3D Gaussian Splatting (3DGS).
Observer bias and inconsistencies in traditional plant phenotyping methods limit the accuracy and reproducibility of fine-grained plant analysis. To overcome these challenges, we developed TomatoMAP, a comprehensive dataset for Solanum lycopersicum using an Internet of Things (IoT) based imaging system with standardized data acquisition protocols. Our dataset contains 64,464 RGB images that capture 12 different plant poses from four camera elevation angles. Each image includes manually annotated bounding boxes for seven regions of interest (ROIs), including leaves, panicle, batch of flowers, batch of fruits, axillary shoot, shoot and whole plant area, along with 50 fine-grained growth stage classifications based on the BBCH scale. Additionally, we provide 3,616 high-resolution image subset with pixel-wise semantic and instance segmentation annotations for fine-grained phenotyping. We validated our dataset using a cascading model deep learning framework combining MobileNetv3 for classification, YOLOv11 for object detection, and MaskRCNN for segmentation. Through AI vs. Human analysis involving five domain experts, we demonstrate that the models trained on our dataset achieve accuracy and speed comparable to the experts. Cohen's Kappa and inter-rater agreement heatmap confirm the reliability of automated fine-grained phenotyping using our approach.
Why it matches plant phenotyping methods植物の多視点画像取得、アノテーション付きデータセット、深層学習による分類・検出・セグメンテーションを中心に開発・検証した、明確な植物フェノタイピング手法研究です。
abstractwe developed TomatoMAP, a comprehensive dataset for Solanum lycopersicum using an Internet of Things (IoT) based imaging system with standardized data acquisition protocols.
Reproduction assets foundThe paper's TomatoMAP dataset (images, annotations) is publicly deposited in e!DAL at IPK with an explicit DOI URL given in the Data Records section.Dataset · publicDataset is deposited in e!DAL (electronic data archive library) of IPK (Leibniz Institute of Plant Genetics
and Crop Plant Research): https://doi.ipk-gatersleben.de/DOI/10bb9f14-ce90-4747-836f-cf61dfb5eea1/Open asset ↗e!DAL · 10bb9f14-ce90-4747-836f-cf61dfb5eea1pdf-page:7 lines:1-73Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 13 Sept 2026
In orchard automation, dense foliage during the canopy season severely occludes tree structures, minimizing visibility to various canopy parts such as trunks and branches, which limits the ability of a machine vision system. However, canopy structure is more open and visible during the dormant season when trees are defoliated. In this work, we present an information fusion framework that integrates multi-seasonal structural data to support robotic and automated crop load management during the entire growing season. The framework combines high-resolution RGB-D imagery from both dormant and canopy periods using YOLOv9-Seg for instance segmentation, Kinect Fusion for 3D reconstruction, and Fast Generalized Iterative Closest Point (Fast GICP) for model alignment. Segmentation outputs from YOLOv9-Seg were used to extract depth-informed masks, which enabled accurate 3D point cloud reconstruction via Kinect Fusion; these reconstructed models from each season were subsequently aligned using Fast GICP to achieve spatially coherent multi-season fusion. The YOLOv9-Seg model, trained on manually annotated images, achieved a mean squared error (MSE) of 0.0047 and segmentation mAP@50 scores up to 0.78 for trunks in dormant season dataset. Kinect Fusion enabled accurate reconstruction of tree geometry, validated with field measurements resulting in root mean square errors (RMSE) of 5.23 mm for trunk diameter, 4.50 mm for branch diameter, and 13.72 mm for branch spacing. Fast GICP achieved precise cross-seasonal registration with a minimum fitness score of 0.00197, allowing integrated, comprehensive tree structure modeling despite heavy occlusions during the growing season. This fused structural representation enables robotic systems to access otherwise obscured architectural information, improving the precision of pruning, thinning, and other automated orchard operations.
Why it matches plant phenotyping methods果樹の幹・枝の形態を3D再構成・季節間融合で推定する画像ベースの表現型取得手法を開発し、実測値で検証しているため、方法が中心である。
abstractIn this work, we present an information fusion framework that integrates multi-seasonal structural data to support robotic and automated crop load management during the entire growing season.
In precision agriculture, one of the most important tasks when exploring crop production is identifying individual plant components. There are several attempts to accomplish this task by the use of traditional 2D imaging, 3D reconstructions, and Convolutional Neural Networks (CNN). However, they have several drawbacks when processing 3D data and identifying individual plant components. Therefore, in this work, we propose a novel Deep Learning architecture to detect components of individual plants on Light Detection and Ranging (LiDAR) 3D Point Cloud (PC) data sets. This architecture is based on the concept of Graph Neural Networks (GNN), and feature enhancing with Principal Component Analysis (PCA). For this, each point is taken as a vertex and by the use of a K-Nearest Neighbors (KNN) layer, the edges are established, thus representing the 3D PC data set. Subsequently, Edge-Conv layers are used to further increase the features of each point. Finally, Graph Attention Networks (GAT) are applied to classify visible phenotypic components of the plant, such as the leaf, stem, and soil. This study demonstrates that our graph-based deep learning approach enhances segmentation accuracy for identifying individual plant components, achieving percentages above 80% in the IoU average, thus outperforming other existing models based on point clouds.
Why it matches plant phenotyping methodsLiDAR 3D点群からトウモロコシの葉・茎などの可視形態構成要素を分割・分類するGNN手法の開発であり、植物表現型の取得・抽出が中心です。
abstractwe propose a novel Deep Learning architecture to detect components of individual plants on Light Detection and Ranging (LiDAR) 3D Point Cloud (PC) data sets
Quantitative descriptions of the complete canopy architecture are essential for accurately evaluating crop photosynthesis and yield performance to guide ideotype design. Although various sensing technologies have been developed for three-dimensional (3D) reconstruction of individual plants and canopies, they failed to obtain an accurate description of canopy architectures due to severe occlusion among complex canopy architectures. We proposed an effective method for 3D reconstruction of complex, dynamic population canopy architecture for rapeseed crops with a novel point cloud completion model. A complete point cloud generation framework was developed for automated annotation of the training dataset by distinguishing surface points from occluded points within canopies. The crop population point cloud completion network (CP-PCN) was then designed with a multi-resolution dynamic graph convolutional encoder (MRDG) and a point pyramid decoder (PPD) to predict occluded points. To further enhance feature extraction, a dynamic graph convolutional feature extractor (DGCFE) module was proposed to capture structural variations over the whole rapeseed growth period. The results demonstrated that CP-PCN achieved chamfer distance (CD) values of 3.35 cm -4.51 cm over four growth stages, outperforming the state-of-the-art transformer-based method (PoinTr). Ablation studies confirmed the effectiveness of the MRDG and DGCFE modules. Moreover, the validation experiment demonstrated that the silique efficiency index developed from CP-PCN improved the overall accuracy of rapeseed yield prediction by 11.2% compared to that of using incomplete point clouds. The CP-PCN pipeline has the potential to be extended to other crops, significantly advancing the quantitatively analysis of in-field population canopy architectures.
Why it matches plant phenotyping methods作物群落キャノピーの3D形態を復元する点群補完法を開発し、既存法との比較、アブレーション、収量予測への有効性検証まで行っており、植物フェノタイピング手法が研究の中心である。
abstractWe proposed an effective method for 3D reconstruction of complex, dynamic population canopy architecture for rapeseed crops with a novel point cloud completion model.
Reproduction assets foundThe paper's availability statement explicitly deposits all source code and test data (rapeseed canopy point cloud completion, CP-PCN) on GitHub at the allowed URL.Code · publicn Wang, Yi Feng,
Mengjie Gong and Guangyu Wu, for their participation in the experiments, and
to the Jiaxing Academy of Agricultural Sciences for their assistance with the
experimental data acquisition.
Availability of supporting data and source code
All source codes and test data involved in this study are available on
GitHub (https://github.com/Ziyue-Guo/RP-PCN.git).
Declaration of Competing Interest
The authors declare that they have no known competing financial interests
or personal relationships that could have appeared to influence the work
reported in this paper.
Contributions
Z. G. designed the study, conducted the experiments, and wrote the
manuscript. Y. S. contributed to the expeOpen asset ↗Ziyue-Guo/RP-PCNpdf-layout-page:42 lines:1-42Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Accurate identification of individual plants from unmanned aerial vehicle (UAV) images is essential for advancing high-throughput phenotyping and supporting data-driven decision-making in plant breeding. This study presents MatchPlant, a modular, graphical user interface-supported, open-source Python pipeline for UAV-based single-plant detection and geospatial trait extraction. MatchPlant enables end-to-end workflows by integrating UAV image processing, user-guided annotation, Convolutional Neural Network model training for object detection, forward projection of bounding boxes onto an orthomosaic, and shapefile generation for spatial phenotypic analysis. In an early-season maize case study, MatchPlant achieved reliable detection performance (validation AP: 89.6%, test AP: 85.9%) and effectively projected bounding boxes, covering 89.8% of manually annotated boxes with 87.5% of projections achieving an Intersection over Union (IoU) greater than 0.5. Trait values extracted from predicted bounding instances showed high agreement with manual annotations (r = 0.87-0.97, IoU >= 0.4). Detection outputs were reused across time points to extract plant height and Normalized Difference Vegetation Index with minimal additional annotation, facilitating efficient temporal phenotyping. By combining modular design, reproducibility, and geospatial precision, MatchPlant offers a scalable framework for UAV-based plant-level analysis with broad applicability in agricultural and environmental monitoring.
Why it matches plant phenotyping methodsUAV画像から個体検出・地理空間的形質抽出を行うオープンソース基盤の開発と性能検証が中心であり、植物形質(草丈・NDVI)を抽出する再利用可能なワークフローを提供している。
abstractThis study presents MatchPlant, a modular, graphical user interface-supported, open-source Python pipeline for UAV-based single-plant detection and geospatial trait extraction.
Reproduction assets foundThe paper's MatchPlant pipeline code is publicly available on GitHub, and the maize case study training dataset and pre-trained model are publicly available on Zenodo.Dataset · publicinistration, Funding acquisition.
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Data availability
The public datasets supporting the case study are available on Zenodo at https://doi.org/10.5281/zenodo.14856123 (accessed on February 14, 2025). The source code and documentation for MatchPlant are available on GitHub at https://github.com/JacobWashburn-USDA/MatchPlant (accessed on February 14, 2025).
Acknowledgments
This research was supported in part by an appointment to the Agricultural Research Service (ARS) Research Participation PrOpen asset ↗Zenodo · 10.5281/zenodo.14856123lines:169-250Plant phenotyping relevance match · UnverifiedarXiv · checked 13 Sept 2026
Field / plotNeRF / 3D Gaussian SplattingFruitCounting2D/3D reconstruction
Accurate 3D fruit counting in orchards is challenging due to heavy occlusion, semantic ambiguity between fruits and surrounding structures, and the high computational cost of volumetric reconstruction. Existing pipelines often rely on multi-view 2D segmentation and dense volumetric sampling, which lead to accumulated fusion errors and slow inference. We introduce FruitLangGS, a language-guided 3D fruit counting framework that reconstructs orchard-scale scenes using an adaptive-density Gaussian Splatting pipeline with radius-aware pruning and tile-based rasterization, enabling scalable 3D representation. During inference, compressed CLIP-aligned semantic vectors embedded in each Gaussian are filtered via a dual-threshold cosine similarity mechanism, retrieving Gaussians relevant to target prompts while suppressing common distractors (e.g., foliage), without requiring retraining or image-space masks. The selected Gaussians are then sampled into dense point clouds and clustered geometrically to estimate fruit instances, remaining robust under severe occlusion and viewpoint variation. Experiments on nine different orchard-scale datasets demonstrate that FruitLangGS consistently outperforms existing pipelines in instance counting recall, avoiding multi-view segmentation fusion errors and achieving up to 99.7% recall on Pfuji-Size_Orch2018 orchard dataset. Ablation studies further confirm that language-conditioned semantic embedding and dual-threshold prompt filtering are essential for suppressing distractors and improving counting accuracy under heavy occlusion. Beyond fruit counting, the same framework enables prompt-driven 3D semantic retrieval without retraining, highlighting the potential of language-guided 3D perception for scalable agricultural scene understanding.
Why it matches plant phenotyping methods果実個数という植物器官形質を推定する3D画像解析手法を開発し、複数データセットで性能検証しているため、植物フェノタイピング手法が研究の中心である。
abstractWe introduce FruitLangGS, a language-guided 3D fruit counting framework
AppleMangoPeachPearPlumField / plotNeRF / 3D Gaussian SplattingRGB / grayscaleFruitWhole plant / canopy / plot / field
FruitNeRF++: A Generalized Multi-Fruit Counting Method Utilizing Contrastive Learning and Neural Radiance Fields We introduce FruitNeRF++, a novel fruit-counting approach that combines contrastive learning with neural radiance fields to count fruits from unstructured input photographs of orchards. Our work is based on FruitNeRF [6], which employs a neural semantic field combined with a fruit-specific clusteringapproach. The requirement for adaptation for each fruit type limits the applicability of the method, and makes it difficult to use in practice. To lift this limitation, we design a shape-agnostic multi-fruit counting framework, that complements the RGB and semantic data with instance masks predicted by a vision foundation model. The masks are used to encode the identity of each fruit as instance embeddings into a neural instance field. By volumetrically sampling the neural fields, we extract apoint cloud embedded with the instance features, which can be clustered in a fruit-agnostic manner to obtain the fruit count. We evaluate our approach using a synthetic dataset containing apples, plums, lemons, pears, peaches, and mangoes, as well as a real-world benchmark apple dataset. Our results demonstrate that FruitNeRF++ is easier to control and compares favorably to other state-of-the-art methods.
Why it matches plant phenotyping methods果実を対象とした画像ベースの汎用カウント手法を開発し、合成および実データで評価しているため、植物形質(果実数)の取得・推定が研究の中心です。
abstractWe introduce FruitNeRF++, a novel fruit-counting approach that combines contrastive learning with neural radiance fields to count fruits from unstructured input photographs of orchards.
NeRF / 3D Gaussian SplattingLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstruction
Recent years have seen substantial improvements in the ability to generate synthetic 3D objects using AI. However, generating complex 3D objects, such as plants, remains a considerable challenge. Current generative 3D models struggle with plant generation compared to general objects, limiting their usability in plant analysis tools, which require fine detail and accurate geometry. We introduce PlantDreamer, a novel approach to 3D synthetic plant generation, which can achieve greater levels of realism for complex plant geometry and textures than available text-to-3D models. To achieve this, our new generation pipeline leverages a depth ControlNet, fine-tuned Low-Rank Adaptation and an adaptable Gaussian culling algorithm, which directly improve textural realism and geometric integrity of generated 3D plant models. Additionally, PlantDreamer enables both purely synthetic plant generation, by leveraging L-System-generated meshes, and the enhancement of real-world plant point clouds by converting them into 3D Gaussian Splats. We evaluate our approach by comparing its outputs with state-of-the-art text-to-3D models, demonstrating that PlantDreamer outperforms existing methods in producing high-fidelity synthetic plants. Our results indicate that our approach not only advances synthetic plant generation, but also facilitates the upgrading of legacy point cloud datasets, making it a valuable tool for 3D phenotyping applications.
Why it matches plant phenotyping methods植物の高精細3Dモデル生成・実点群のGaussian Splat変換を開発し、3Dフェノタイピングへの利用を明示的に評価しているため、手法が中心である。
abstractWe introduce PlantDreamer, a novel approach to 3D synthetic plant generation
Deep learning has transformed computer vision for precision agriculture, yet apple orchard monitoring remains limited by dataset constraints. The lack of diverse, realistic datasets and the difficulty of annotating dense, heterogeneous scenes. Existing datasets overlook different growth stages and stereo imagery, both essential for realistic 3D modeling of orchards and tasks like fruit localization, yield estimation, and structural analysis. To address these gaps, we present AppleGrowthVision, a large-scale dataset comprising two subsets. The first includes 9,317 high resolution stereo images collected from a farm in Brandenburg (Germany), covering six agriculturally validated growth stages over a full growth cycle. The second subset consists of 1,125 densely annotated images from the same farm in Brandenburg and one in Pillnitz (Germany), containing a total of 31,084 apple labels. AppleGrowthVision provides stereo-image data with agriculturally validated growth stages, enabling precise phenological analysis and 3D reconstructions. Extending MinneApple with our data improves YOLOv8 performance by 7.69 % in terms of F1-score, while adding it to MinneApple and MAD boosts Faster R-CNN F1-score by 31.06 %. Additionally, six BBCH stages were predicted with over 95 % accuracy using VGG16, ResNet152, DenseNet201, and MobileNetv2. AppleGrowthVision bridges the gap between agricultural science and computer vision, by enabling the development of robust models for fruit detection, growth modeling, and 3D analysis in precision agriculture. Future work includes improving annotation, enhancing 3D reconstruction, and extending multimodal analysis across all growth stages.
Why it matches plant phenotyping methodsリンゴの生育段階・果実・樹体構造を対象とする大規模ステレオ画像データセットを構築し、果実検出、フェノロジー分析、3D再構成モデルの評価に用いており、表現型取得・解析基盤が中心である。
titleAppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards
Field / plotPhotogrammetry / SfM / MVSLiDAR / point cloudRGB / grayscaleStem / branchMorphology / geometry measurement2D/3D reconstructionSegmentationVisualization / data managementArchitecture / morphology / geometry
Forest inventories rely on accurate measurements of the diameter at breast height (DBH) for ecological monitoring, resource management, and carbon accounting. While LiDAR-based techniques can achieve centimeter-level precision, they are cost-prohibitive and operationally complex. We present a low-cost alternative that only needs a consumer-grade 360 video camera. Our semi-automated pipeline comprises of (i) a dense point cloud reconstruction using Structure from Motion (SfM) photogrammetry software called Agisoft Metashape, (ii) semantic trunk segmentation by projecting Grounded Segment Anything (SAM) masks onto the 3D cloud, and (iii) a robust RANSAC-based technique to estimate cross section shape and DBH. We introduce an interactive visualization tool for inspecting segmented trees and their estimated DBH. On 61 acquisitions of 43 trees under a variety of conditions, our method attains median absolute relative errors of 5-9% with respect to "ground-truth" manual measurements. This is only 2-4% higher than LiDAR-based estimates, while employing a single 360 camera that costs orders of magnitude less, requires minimal setup, and is widely available.
Why it matches plant phenotyping methodsRGB映像から樹木のDBHという明示的な植物形態形質を推定する半自動パイプラインを開発し、手動測定およびLiDARと比較して精度検証しているため、植物フェノタイピング手法が中心である。
abstractWe present a low-cost alternative that only needs a consumer-grade 360 video camera.
As the agricultural workforce declines and labor costs rise, robotic yield estimation has become increasingly important. While unmanned ground vehicles (UGVs) are commonly used for indoor farm monitoring, their deployment in greenhouses is often constrained by infrastructure limitations, sensor placement challenges, and operational inefficiencies. To address these issues, we develop a lightweight unmanned aerial vehicle (UAV) equipped with an RGB-D camera, a 3D LiDAR, and an IMU sensor. The UAV employs a LiDAR-inertial odometry algorithm for precise navigation in GNSS-denied environments and utilizes a 3D multi-object tracking algorithm to estimate the count and weight of cherry tomatoes. We evaluate the system using two dataset: one from a harvesting row and another from a growing row. In the harvesting-row dataset, the proposed system achieves 94.4\% counting accuracy and 87.5\% weight estimation accuracy within a 13.2-meter flight completed in 10.5 seconds. For the growing-row dataset, which consists of occluded unripened fruits, we qualitatively analyze tracking performance and highlight future research directions for improving perception in greenhouse with strong occlusions. Our findings demonstrate the potential of UAVs for efficient robotic yield estimation in commercial greenhouses.
Why it matches plant phenotyping methodsUAV上のRGB-D・LiDAR・追跡アルゴリズムにより、トマト果実数と重量を推定する手法を開発・評価しており、植物形質取得が研究の中心です。
abstractwe develop a lightweight unmanned aerial vehicle (UAV) equipped with an RGB-D camera, a 3D LiDAR, and an IMU sensor.
Plant phenotyping increasingly relies on (semi-)automated image-based analysis workflows to improve its accuracy and scalability. However, many existing solutions remain overly complex, difficult to reimplement and maintain, and pose high barriers for users without substantial computational expertise. To address these challenges, we introduce PhenoAssistant: a pioneering AI-driven system that streamlines plant phenotyping via intuitive natural language interaction. PhenoAssistant leverages a large language model to orchestrate a curated toolkit supporting tasks including automated phenotype extraction, data visualisation and automated model training. We validate PhenoAssistant through several representative case studies and a set of evaluation tasks. By significantly lowering technical hurdles, PhenoAssistant underscores the promise of AI-driven methodologies to democratising AI adoption in plant biology.
Why it matches plant phenotyping methods植物フェノタイピングの画像解析ワークフローを自然言語で自動化するシステムを開発し、ケーススタディと評価タスクで検証しているため、方法が中心である。
abstractwe introduce PhenoAssistant: a pioneering AI-driven system that streamlines plant phenotyping via intuitive natural language interaction.
Reproduction assets foundThe paper's authors release PhenoAssistant's code, chat logs, and evaluation results on GitHub, and the winter wheat nutrient-deficiency dataset used in Case Study 3 is publicly available on CodaLab. Case Study 1 demonstration data is request-only (Phenotiki), and Case Study 2 data is on Zenodo, which is not among the审Dataset · publicData for demonstrating Case Study 3 are publicly
available at https://codalab.lisn.upsaclay.fr/competitions/13833.Open asset ↗pdf-page:13 lines:1-47Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 15 Sept 2026
We present an open-source, low-cost photogrammetry system for 3D plant modeling and phenotyping. The system uses a structure-from-motion approach to reconstruct 3D representations of the plants via point clouds. Using wheat as an example, we demonstrate how various phenotypic traits can be computed easily from the point clouds. These include standard measurements such as plant height and radius, as well as features that would be more cumbersome to measure by hand, such as leaf angles and convex hull. We further demonstrate the utility of the system through the investigation of specific metrics that may yield objective classifications of erectophile versus planophile wheat canopy architectures.
Why it matches plant phenotyping methods低コストの3Dフォトグラメトリによる植物形状再構成と、点群からの表現型形質推定を中心に開発・実証しているため。
abstractWe present an open-source, low-cost photogrammetry system for 3D plant modeling and phenotyping.
ArabidopsisTomatoLeafRootSeed / grainSegmentationGrowth / time-series analysisTrackingGrowth / development / phenologyRoot system architecture
Plant developmental plasticity, particularly in root system architecture, is fundamental to understanding adaptability and agricultural sustainability. ChronoRoot 2.0 builds upon established low-cost hardware while significantly enhancing software capabilities and usability. The system employs nnUNet architecture for multi-class segmentation, demonstrating significant accuracy improvements while simultaneously tracking six distinct plant structures encompassing root, shoot, and seed components: main root, lateral roots, seed, hypocotyl, leaves, and petiole. This architecture enables easy retraining and incorporation of additional training data without requiring machine learning expertise. The platform introduces dual specialized graphical interfaces: a Standard Interface for detailed architectural analysis with novel gravitropic response parameters, and a Screening Interface enabling high-throughput analysis of multiple plants through automated tracking. Functional Principal Component Analysis integration enables discovery of novel phenotypic parameters through temporal pattern comparison. We demonstrate multi-species analysis, with Arabidopsis thaliana and Solanum lycopersicum, both morphologically distinct plant species. Three use cases in Arabidopsis thaliana and validation with tomato seedlings demonstrate enhanced capabilities: circadian growth pattern characterization, gravitropic response analysis in transgenic plants, and high-throughput etiolation screening across multiple genotypes.ChronoRoot 2.0 maintains the low-cost, modular hardware advantages of its predecessor while dramatically improving accessibility through intuitive graphical interfaces and expanded analytical capabilities. The open-source platform makes sophisticated temporal plant phenotyping more accessible to researchers without computational expertise.
Why it matches plant phenotyping methods植物の時系列画像から根・地上部・種子などの形態形質を抽出・追跡するオープンなAI基盤を開発し、精度向上、再学習、GUI、高スループット解析、検証まで扱っており、フェノタイピング手法が研究の中心です。
titleChronoRoot 2.0: An Open AI-Powered Platform for 2D Temporal Plant Phenotyping
Reproduction assets foundThe paper explicitly releases its full analysis source code (GitHub), the annotated infrared image dataset with multiclass segmentation masks (HuggingFace), demo phenotype video datasets, and a Docker image — all paper-specific, public, and actionable.Code · publicThe complete source code of ChronoRoot 2.0, including the implementation of all analysis methods described in this paper, is freely available under the GNU General Public License v3.0 at https://github.com/ChronoRoot/ChronoRoot2Open asset ↗ChronoRoot/ChronoRoot2lines:491-523Dataset · publicThe annotated image dataset used for training and validation contains 911 infrared images of Arabidopsis thaliana seedlings and 480 images of tomato with expert annotations for multiclass segmentation. This dataset is publicly available without restrictions at https://huggingface.co/datasets/ngaggion/ChronoRoot2Open asset ↗ngaggion/ChronoRoot2lines:491-523Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 13 Sept 2026
Monitoring flowers over time is essential for precision robotic pollination in agriculture. To accomplish this, a continuous spatial-temporal observation of plant growth can be done using stationary RGB-D cameras. However, image registration becomes a serious challenge due to changes in the visual appearance of the plant caused by the pollination process and occlusions from growth and camera angles. Plants flower in a manner that produces distinct clusters on branches. This paper presents a method for matching flower clusters using descriptors generated from RGB-D data and considers allowing for spatial uncertainty within the cluster. The proposed approach leverages the Unscented Transform to efficiently estimate plant descriptor uncertainty tolerances, enabling a robust image-registration process despite temporal changes. The Unscented Transform is used to handle the nonlinear transformations by propagating the uncertainty of flower positions to determine the variations in the descriptor domain. A Monte Carlo simulation is used to validate the Unscented Transform results, confirming our method's effectiveness for flower cluster matching. Therefore, it can facilitate improved robotics pollination in dynamic environments.
Why it matches plant phenotyping methodsRGB-D画像から花房を時系列追跡・マッチングする画像登録手法が研究の中心で、植物成長の観測に直接用いられるため、植物フェノタイピング手法として含める。
abstractThis paper presents a method for matching flower clusters using descriptors generated from RGB-D data and considers allowing for spatial uncertainty within the cluster.
Accurate estimation of total leaf area (TLA) is crucial for evaluating plant growth, photosynthetic activity, and transpiration. However, it remains challenging for bushy plants like dwarf tomatoes due to their complex canopies. Traditional methods are often labor-intensive, damaging to plants, or limited in capturing canopy complexity. This study evaluated a non-destructive method combining sequential 3D reconstructions from RGB images and machine learning to estimate TLA for three dwarf tomato cultivars: Mohamed, Hahms Gelbe Topftomate, and Red Robin -- grown under controlled greenhouse conditions. Two experiments (spring-summer and autumn-winter) included 73 plants, yielding 418 TLA measurements via an "onion" approach. High-resolution videos were recorded, and 500 frames per plant were used for 3D reconstruction. Point clouds were processed using four algorithms (Alpha Shape, Marching Cubes, Poisson's, Ball Pivoting), and meshes were evaluated with seven regression models: Multivariable Linear Regression, Lasso Regression, Ridge Regression, Elastic Net Regression, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron. The Alpha Shape reconstruction ($α= 3$) with Extreme Gradient Boosting achieved the best performance ($R^2 = 0.80$, $MAE = 489 cm^2$). Cross-experiment validation showed robust results ($R^2 = 0.56$, $MAE = 579 cm^2$). Feature importance analysis identified height, width, and surface area as key predictors. This scalable, automated TLA estimation method is suited for urban farming and precision agriculture, offering applications in automated pruning, resource efficiency, and sustainable food production. The approach demonstrated robustness across variable environmental conditions and canopy structures.
Why it matches plant phenotyping methodsRGB画像からの3D再構成と機械学習により、植物形質である総葉面積を非破壊・自動推定する方法を開発・検証しており、フェノタイピング手法が研究の中心です。
abstractThis study evaluated a non-destructive method combining sequential 3D reconstructions from RGB images and machine learning to estimate TLA for three dwarf tomato cultivars
Crop yield estimation is a relevant problem in agriculture, because an accurate yield estimate can support farmers' decisions on harvesting or precision intervention. Robots can help to automate this process. To do so, they need to be able to perceive the surrounding environment to identify target objects such as trees and plants. In this paper, we introduce a novel approach to address the problem of hierarchical panoptic segmentation of apple orchards on 3D data from different sensors. Our approach is able to simultaneously provide semantic segmentation, instance segmentation of trunks and fruits, and instance segmentation of trees (a trunk with its fruits). This allows us to identify relevant information such as individual plants, fruits, and trunks, and capture the relationship among them, such as precisely estimate the number of fruits associated to each tree in an orchard. To efficiently evaluate our approach for hierarchical panoptic segmentation, we provide a dataset designed specifically for this task. Our dataset is recorded in Bonn, Germany, in a real apple orchard with a variety of sensors, spanning from a terrestrial laser scanner to a RGB-D camera mounted on different robots platforms. The experiments show that our approach surpasses state-of-the-art approaches in 3D panoptic segmentation in the agricultural domain, while also providing full hierarchical panoptic segmentation. Our dataset is publicly available at https://www.ipb.uni-bonn.de/data/hops/. The open-source implementation of our approach is available at https://github.com/PRBonn/hapt3D.
Why it matches plant phenotyping methodsリンゴ樹・果実・幹を3Dセグメンテーションし、樹ごとの果実数を推定する手法と専用データセットを中心に開発・評価しており、植物の器官形態・収量関連形質の取得に該当する。
abstractwe introduce a novel approach to address the problem of hierarchical panoptic segmentation of apple orchards on 3D data from different sensors.
Reproduction assets foundThe paper introduces the HOPS dataset of annotated 3D apple orchard point clouds (TLS, UAV, UGV, SfM) for hierarchical panoptic segmentation, publicly available at the authors' IPB Bonn page, and releases the open-source implementation (hapt3D) on GitHub. Both are paper-specific, public, and actionable.Code · publicThe open-source implementation of our approach is available at https://github.com/PRBonn/hapt3D .Open asset ↗PRBonn/hapt3Dlines:1-59Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 15 Sept 2026
Estimation of a single leaf area can be a measure of crop growth and a phenotypic trait to breed new varieties. It has also been used to measure leaf area index and total leaf area. Some studies have used hand-held cameras, image processing 3D reconstruction and unsupervised learning-based methods to estimate the leaf area in plant images. Deep learning works well for object detection and segmentation tasks; however, direct area estimation of objects has not been explored. This work investigates deep learning-based leaf area estimation, for RGBD images taken using a mobile camera setup in real-world scenarios. A dataset for attached leaves captured with a top angle view and a dataset for detached single leaves were collected for model development and testing. First, image processing-based area estimation was tested on manually segmented leaves. Then a Mask R-CNN-based model was investigated, and modified to accept RGBD images and to estimate the leaf area. The detached-leaf data set was then mixed with the attached-leaf plant data set to estimate the single leaf area for plant images, and another network design with two backbones was proposed: one for segmentation and the other for area estimation. Instead of trying all possibilities or random values, an agile approach was used in hyperparameter tuning. The final model was cross-validated with 5-folds and tested with two unseen datasets: detached and attached leaves. The F1 score with 90% IoA for segmentation result on unseen detached-leaf data was 1.0, while R-squared of area estimation was 0.81. For unseen plant data segmentation, the F1 score with 90% IoA was 0.59, while the R-squared score was 0.57. The research suggests using attached leaves with ground truth area to improve the results.
Why it matches plant phenotyping methodsRGBD画像と深層学習を用いて葉面積という植物形質を推定する手法を開発し、交差検証と未知データで性能評価しており、フェノタイピング手法が中心である。
abstractThis work investigates deep learning-based leaf area estimation, for RGBD images taken using a mobile camera setup in real-world scenarios.
High-density planting is a widely adopted strategy to enhance maize productivity, yet it introduces challenges such as increased interplant competition and shading, which can limit light capture and overall yield potential. In response, some maize plants naturally reorient their canopies to optimize light capture, a process known as canopy reorientation. Understanding this adaptive response and its impact on light capture is crucial for maximizing agricultural yield potential. This study introduces an end-to-end framework that integrates realistic 3D reconstructions of field-grown maize with photosynthetically active radiation (PAR) modeling to assess the effects of phyllotaxy and planting density on light interception. In particular, using 3D point clouds derived from field data, virtual fields for a diverse set of maize genotypes were constructed and validated against field PAR measurements. Using this framework, we present detailed analyses of the impact of canopy orientations, plant and row spacings, and planting row directions on PAR interception throughout a typical growing season. Our findings highlight significant variations in light interception efficiency across different planting densities and canopy orientations. By elucidating the relationship between canopy architecture and light capture, this study offers valuable guidance for optimizing maize breeding and cultivation strategies across diverse agricultural settings.
Why it matches plant phenotyping methods圃場トウモロコシの3D再構成とPARモデルを統合し、圃場測定で検証した再利用可能な表現型取得・解析フレームワークが研究の中心である。
abstractThis study introduces an end-to-end framework that integrates realistic 3D reconstructions of field-grown maize with photosynthetically active radiation (PAR) modeling to assess the effects of phyllotaxy and planting density on light interception.
Understanding plant growth dynamics is essential for applications in agriculture and plant phenotyping. We present the Growth Modelling (GroMo) challenge, which is designed for two primary tasks: (1) plant age prediction and (2) leaf count estimation, both essential for crop monitoring and precision agriculture. For this challenge, we introduce GroMo25, a dataset with images of four crops: radish, okra, wheat, and mustard. Each crop consists of multiple plants (p1, p2, ..., pn) captured over different days (d1, d2, ..., dm) and categorized into five levels (L1, L2, L3, L4, L5). Each plant is captured from 24 different angles with a 15-degree gap between images. Participants are required to perform both tasks for all four crops with these multiview images. We proposed a Multiview Vision Transformer (MVVT) model for the GroMo challenge and evaluated the crop-wise performance on GroMo25. MVVT reports an average MAE of 7.74 for age prediction and an MAE of 5.52 for leaf count. The GroMo Challenge aims to advance plant phenotyping research by encouraging innovative solutions for tracking and predicting plant growth. The GitHub repository is publicly available at https://github.com/mriglab/GroMo-Plant-Growth-Modeling-with-Multiview-Images.
Why it matches plant phenotyping methods植物のマルチビュー画像から葉数と植物齢を推定するデータセット・ベンチマークおよびモデルを提示しており、植物表現型取得・推定が中心である。
abstractWe present the Growth Modelling (GroMo) challenge, which is designed for two primary tasks: (1) plant age prediction and (2) leaf count estimation
NeRF / 3D Gaussian SplattingFlowerPose / keypoint estimation
This study presents Flower Pose Estimation (FloPE), a real-time flower pose estimation framework for computationally constrained robotic pollination systems. Robotic pollination has been proposed to supplement natural pollination to ensure global food security due to the decreased population of natural pollinators. However, flower pose estimation for pollination is challenging due to natural variability, flower clusters, and high accuracy demands due to the flowers' fragility when pollinating. This method leverages 3D Gaussian Splatting to generate photorealistic synthetic datasets with precise pose annotations, enabling effective knowledge distillation from a high-capacity teacher model to a lightweight student model for efficient inference. The approach was evaluated on both single and multi-arm robotic platforms, achieving a mean pose estimation error of 0.6 cm and 19.14 degrees within a low computational cost. Our experiments validate the effectiveness of FloPE, achieving up to 78.75% pollination success rate and outperforming prior robotic pollination techniques.
Why it matches plant phenotyping methods花の姿勢という植物器官の形態的状態を推定する画像・計算手法を開発し、ロボット実機で性能検証しており、フェノタイピング手法が中心である。
abstractThis study presents Flower Pose Estimation (FloPE), a real-time flower pose estimation framework for computationally constrained robotic pollination systems.
Aerial / UAVLiDAR / point cloudMultispectral / hyperspectralLeafStem / branchSegmentation
Point clouds captured with laser scanning systems from forest environments can be utilized in a wide variety of applications within forestry and plant ecology, such as the estimation of tree stem attributes, leaf angle distribution, and above-ground biomass. However, effectively utilizing the data in such tasks requires the semantic segmentation of the data into wood and foliage points, also known as leaf-wood separation. The traditional approach to leaf-wood separation has been geometry- and radiometry-based unsupervised algorithms, which tend to perform poorly on data captured with airborne laser scanning (ALS) systems, even with a high point density. While recent machine and deep learning approaches achieve great results even on sparse point clouds, they require manually labeled training data, which is often extremely laborious to produce. Multispectral (MS) information has been demonstrated to have potential for improving the accuracy of leaf-wood separation, but quantitative assessment of its effects has been lacking. This study proposes a fully unsupervised deep learning method, GrowSP-ForMS, which is specifically designed for leaf-wood separation of high-density MS ALS point clouds and based on the GrowSP architecture. GrowSP-ForMS achieved a mean accuracy of 84.3% and a mean intersection over union (mIoU) of 69.6% on our MS test set, outperforming the unsupervised reference methods by a significant margin. When compared to supervised deep learning methods, our model performed similarly to the slightly older PointNet architecture but was outclassed by more recent approaches. Finally, two ablation studies were conducted, which demonstrated that our proposed changes increased the test set mIoU of GrowSP-ForMS by 29.4 percentage points (pp) in comparison to the original GrowSP model and that utilizing MS data improved the mIoU by 5.6 pp from the monospectral case.
Why it matches plant phenotyping methods森林点云から木部・葉部を分離する深層学習手法を開発し、精度比較とアブレーション検証を行っており、植物構造・属性抽出の中核的方法論である。
abstractThis study proposes a fully unsupervised deep learning method, GrowSP-ForMS, which is specifically designed for leaf-wood separation of high-density MS ALS point clouds
This study introduces a robust framework for generating procedural 3D models of maize (Zea mays) plants from LiDAR point cloud data, offering a scalable alternative to traditional field-based phenotyping. Our framework leverages Non-Uniform Rational B-Spline (NURBS) surfaces to model the leaves of maize plants, combining Particle Swarm Optimization (PSO) for an initial approximation of the surface and a differentiable programming framework for precise refinement of the surface to fit the point cloud data. In the first optimization phase, PSO generates an approximate NURBS surface by optimizing its control points, aligning the surface with the LiDAR data, and providing a reliable starting point for refinement. The second phase uses NURBS-Diff, a differentiable programming framework, to enhance the accuracy of the initial fit by refining the surface geometry and capturing intricate leaf details. Our results demonstrate that, while PSO establishes a robust initial fit, the integration of differentiable NURBS significantly improves the overall quality and fidelity of the reconstructed surface. This hierarchical optimization strategy enables accurate 3D reconstruction of maize leaves across diverse genotypes, facilitating the subsequent extraction of complex traits like phyllotaxy. We demonstrate our approach on diverse genotypes of field-grown maize plants. All our codes are open-source to democratize these phenotyping approaches.
Why it matches plant phenotyping methodsLiDAR点群からトウモロコシ葉の3D形状を再構成し、表現型形質抽出に用いる計算手法を開発・実証しており、植物フェノタイピング手法が中心である。
abstractThis study introduces a robust framework for generating procedural 3D models of maize (Zea mays) plants from LiDAR point cloud data, offering a scalable alternative to traditional field-based phenotyping.
Precision agriculture leverages data and machine learning so that farmers can monitor their crops and target interventions precisely. This enables the precision application of herbicide only to weeds, or the precision application of fertilizer only to undernourished crops, rather than to the entire field. The approach promises to maximize yields while minimizing resource use and harm to the surrounding environment. To this end, we propose a hierarchical panoptic segmentation method that simultaneously determines leaf count (as an identifier of plant growth)and locates weeds within an image. In particular, our approach aims to improve the segmentation of smaller instances like the leaves and weeds by incorporating focal loss and boundary loss. Not only does this result in competitive performance, achieving a PQ+ of 81.89 on the standard training set, but we also demonstrate we can improve leaf-counting accuracy with our method. The code is available at https://github.com/madeleinedarbyshire/HierarchicalMask2Former.
Why it matches plant phenotyping methods植物・葉の階層的パノプティックセグメンテーション法を開発し、葉数という植物成長形質の推定精度を評価しているため、フェノタイピング手法が中心である。
abstractwe propose a hierarchical panoptic segmentation method that simultaneously determines leaf count (as an identifier of plant growth)and locates weeds within an image.
Tree height estimation serves as an important proxy for biomass estimation in ecological and forestry applications. While traditional methods such as photogrammetry and Light Detection and Ranging (LiDAR) offer accurate height measurements, their application on a global scale is often cost-prohibitive and logistically challenging. In contrast, remote sensing techniques, particularly 3D tomographic reconstruction from Synthetic Aperture Radar (SAR) imagery, provide a scalable solution for global height estimation. SAR images have been used in earth observation contexts due to their ability to work in all weathers, unobscured by clouds. In this study, we use deep learning to estimate forest canopy height directly from 2D Single Look Complex (SLC) images, a derivative of SAR. Our method attempts to bypass traditional tomographic signal processing, potentially reducing latency from SAR capture to end product. We also quantify the impact of varying numbers of SLC images on height estimation accuracy, aiming to inform future satellite operations and optimize data collection strategies. Compared to full tomographic processing combined with deep learning, our minimal method (partial processing + deep learning) falls short, with an error 16-21\% higher, highlighting the continuing relevance of geometric signal processing.
Why it matches plant phenotyping methodsSAR画像と深層学習による森林キャノピー高推定手法の開発・比較・精度評価が中心であり、植物群落の明示的な形態形質を測定している。
abstractIn this study, we use deep learning to estimate forest canopy height directly from 2D Single Look Complex (SLC) images, a derivative of SAR.
The ability to automatically build 3D digital twins of plants from images has countless applications in agriculture, environmental science, robotics, and other fields. However, current 3D reconstruction methods fail to recover complete shapes of plants due to heavy occlusion and complex geometries. In this work, we present a novel method for 3D modeling of agricultural crops based on optimizing a parametric model of plant morphology via inverse procedural modeling. Our method first estimates depth maps by fitting a neural radiance field and then optimizes a specialized loss to estimate morphological parameters that result in consistent depth renderings. The resulting 3D model is complete and biologically plausible. We validate our method on a dataset of real images of agricultural fields, and demonstrate that the reconstructed canopies can be used for a variety of monitoring and simulation applications.
Why it matches plant phenotyping methods画像から作物の完全な3D形状・キャノピー構造を再構成する手法を開発し、実画像データセットで検証しているため、植物表現型取得が中心的です。
abstractwe present a novel method for 3D modeling of agricultural crops based on optimizing a parametric model of plant morphology via inverse procedural modeling.
Forest mapping provides critical observational data needed to understand the dynamics of forest environments. Notably, tree diameter at breast height (DBH) is a metric used to estimate forest biomass and carbon dioxide sequestration. Manual methods of forest mapping are labor intensive and time consuming, a bottleneck for large-scale mapping efforts. Automated mapping relies on acquiring dense forest reconstructions, typically in the form of point clouds. Terrestrial laser scanning (TLS) and mobile laser scanning (MLS) generate point clouds using expensive LiDAR sensing, and have been used successfully to estimate tree diameter. Neural radiance fields (NeRFs) are an emergent technology enabling photorealistic, vision-based reconstruction by training a neural network on a sparse set of input views. In this paper, we present a comparison of MLS and NeRF forest reconstructions for the purpose of trunk diameter estimation in a mixed-evergreen Redwood forest. In addition, we propose an improved DBH-estimation method using convex-hull modeling. Using this approach, we achieved 1.68 cm RMSE, which consistently outperformed standard cylinder modeling approaches. Our code contributions and forest datasets are freely available at https://github.com/harelab-ucsc/RedwoodNeRF.
Why it matches plant phenotyping methodsNeRFおよびMLSによる森林再構成から樹木DBHを推定し、凸包モデルによる推定法を提案・比較検証しているため、植物形質取得手法が中心です。
abstractIn this paper, we present a comparison of MLS and NeRF forest reconstructions for the purpose of trunk diameter estimation in a mixed-evergreen Redwood forest.
Reproduction assets foundThe authors explicitly state their code contributions and forest datasets (SLAM and NeRF reconstructions used for DBH estimation) are freely available in a public GitHub repository. The other URLs are a cited third-party tool (NeRFCapture) and a background reference (USDA aerial survey), neither of which is a paper-ownDataset · publicOur code contributions and forest datasets are freely available at https://github.com/harelab-ucsc/RedwoodNeRF .Open asset ↗harelab-ucsc/RedwoodNeRFlines:1-52Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 15 Sept 2026
Field / plotNeRF / 3D Gaussian SplattingPhotogrammetry / SfM / MVSLiDAR / point cloudStem / branchWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionVisualization / data managementArchitecture / morphology / geometry
Accurate and efficient 3D reconstruction of trees is crucial for forest resource assessments and management. Close-Range Photogrammetry (CRP) is commonly used for reconstructing forest scenes but faces challenges like low efficiency and poor quality. Recently, Novel View Synthesis (NVS) technologies, including Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have shown promise for 3D plant reconstruction with limited images. However, existing research mainly focuses on small plants in orchards or individual trees, leaving uncertainty regarding their application in larger, complex forest stands. In this study, we collected sequential images of forest plots with varying complexity and performed dense reconstruction using NeRF and 3DGS. The resulting point clouds were compared with those from photogrammetry and laser scanning. Results indicate that NVS methods significantly enhance reconstruction efficiency. Photogrammetry struggles with complex stands, leading to point clouds with excessive canopy noise and incorrectly reconstructed trees, such as duplicated trunks. NeRF, while better for canopy regions, may produce errors in ground areas with limited views. The 3DGS method generates sparser point clouds, particularly in trunk areas, affecting diameter at breast height (DBH) accuracy. All three methods can extract tree height information, with NeRF yielding the highest accuracy; however, photogrammetry remains superior for DBH accuracy. These findings suggest that NVS methods have significant potential for 3D reconstruction of forest stands, offering valuable support for complex forest resource inventory and visualization tasks.
Why it matches plant phenotyping methods森林スタンドの3D再構成手法を比較・評価し、樹高やDBHという個体樹木形質の抽出精度を検証しているため、植物フェノタイピング手法が中心です。
abstractperformed dense reconstruction using NeRF and 3DGS. The resulting point clouds were compared with those from photogrammetry and laser scanning.
The 3D reconstruction of plants is challenging due to their complex shape causing many occlusions. Next-Best-View (NBV) methods address this by iteratively selecting new viewpoints to maximize information gain (IG). Deep-learning-based NBV (DL-NBV) methods demonstrate higher computational efficiency over classic voxel-based NBV approaches but current methods require extensive training using ground-truth plant models, making them impractical for real-world plants. These methods, moreover, rely on offline training with pre-collected data, limiting adaptability in changing agricultural environments. This paper proposes a self-supervised learning-based NBV method (SSL-NBV) that uses a deep neural network to predict the IG for candidate viewpoints. The method allows the robot to gather its own training data during task execution by comparing new 3D sensor data to the earlier gathered data and by employing weakly-supervised learning and experience replay for efficient online learning. Comprehensive evaluations were conducted in simulation and real-world environments using cross-validation. The results showed that SSL-NBV required fewer views for plant reconstruction than non-NBV methods and was over 800 times faster than a voxel-based method. SSL-NBV reduced training annotations by over 90% compared to a baseline DL-NBV. Furthermore, SSL-NBV could adapt to novel scenarios through online fine-tuning. Also using real plants, the results showed that the proposed method can learn to effectively plan new viewpoints for 3D plant reconstruction. Most importantly, SSL-NBV automated the entire network training and uses continuous online learning, allowing it to operate in changing agricultural environments.
Why it matches plant phenotyping methods植物の3D再構成に向けたロボット視点計画手法を開発し、シミュレーションと実植物で性能評価しているため、表現型取得の中核手法に該当する。
abstractThis paper proposes a self-supervised learning-based NBV method (SSL-NBV) that uses a deep neural network to predict the IG for candidate viewpoints.
An early, non-invasive, and on-site detection of nutrient deficiencies is critical to enable timely actions to prevent major losses of crops caused by lack of nutrients. While acquiring labeled data is very expensive, collecting images from multiple views of a crop is straightforward. Despite its relevance for practical applications, unsupervised domain adaptation where multiple views are available for the labeled source domain as well as the unlabeled target domain is an unexplored research area. In this work, we thus propose an approach that leverages multiple camera views in the source and target domain for unsupervised domain adaptation. We evaluate the proposed approach on two nutrient deficiency datasets. The proposed method achieves state-of-the-art results on both datasets compared to other unsupervised domain adaptation methods. The dataset and source code are available at https://github.com/jh-yi/MV-Match.
Why it matches plant phenotyping methods植物の栄養欠乏状態を画像から推定するマルチビュー・ドメイン適応手法を提案し、2つのデータセットで評価しており、表現型取得・推定手法が中心です。
abstractwe thus propose an approach that leverages multiple camera views in the source and target domain for unsupervised domain adaptation.
Reproduction assets foundThe paper's authors explicitly state that the MiPlo nutrient-deficiency image datasets and the MV-Match source code are publicly available at the authors' GitHub repository, which matches an allowed URL.Dataset · publicThe dataset and source code are available at https://github.com/jh-yi/MV-Match .Open asset ↗jh-yi/MV-Matchlines:1-71Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
AppleCitrusMangoPeachPearPlumField / plotLiDAR / point cloudRGB / grayscaleFruit
We introduce FruitNeRF, a unified novel fruit counting framework that leverages state-of-the-art view synthesis methods to count any fruit type directly in 3D. Our framework takes an unordered set of posed images captured by a monocular camera and segments fruit in each image. To make our system independent of the fruit type, we employ a foundation model that generates binary segmentation masks for any fruit. Utilizing both modalities, RGB and semantic, we train a semantic neural radiance field. Through uniform volume sampling of the implicit Fruit Field, we obtain fruit-only point clouds. By applying cascaded clustering on the extracted point cloud, our approach achieves precise fruit count.The use of neural radiance fields provides significant advantages over conventional methods such as object tracking or optical flow, as the counting itself is lifted into 3D. Our method prevents double counting fruit and avoids counting irrelevant fruit.We evaluate our methodology using both real-world and synthetic datasets. The real-world dataset consists of three apple trees with manually counted ground truths, a benchmark apple dataset with one row and ground truth fruit location, while the synthetic dataset comprises various fruit types including apple, plum, lemon, pear, peach, and mango.Additionally, we assess the performance of fruit counting using the foundation model compared to a U-Net.
Why it matches plant phenotyping methods果実を対象に、画像・NeRF・点群クラスタリングを組み合わせて3D果実数を推定する手法を開発し、実データおよび合成データで評価しているため、植物表現型取得法が中心である。
abstractWe introduce FruitNeRF, a unified novel fruit counting framework that leverages state-of-the-art view synthesis methods to count any fruit type directly in 3D.
Reproduction assets foundThe paper's real-world apple tree image dataset with manual ground-truth counts and synthetic Blender fruit tree data are publicly released via the project website, and the FruitNeRF analysis code is open-source on GitHub. The Zenodo DOI refers to the third-party BlenderNeRF plugin (cited tool), not a paper-specific.Dataset · publicThe data has been made publicly available, and visualizations can be accessed on the project website.Open asset ↗lines:183-221Code · publicFruitNeRF code: https://github.com/meyerls/FruitNeRF has been made open-source.Open asset ↗meyerls/FruitNeRFlines:74-108Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Potato yield is an important metric for farmers to further optimize their cultivation practices. Potato yield can be estimated on a harvester using an RGB-D camera that can estimate the three-dimensional (3D) volume of individual potato tubers. A challenge, however, is that the 3D shape derived from RGB-D images is only partially completed, underestimating the actual volume. To address this issue, we developed a 3D shape completion network, called CoRe++, which can complete the 3D shape from RGB-D images. CoRe++ is a deep learning network that consists of a convolutional encoder and a decoder. The encoder compresses RGB-D images into latent vectors that are used by the decoder to complete the 3D shape using the deep signed distance field network (DeepSDF). To evaluate our CoRe++ network, we collected partial and complete 3D point clouds of 339 potato tubers on an operational harvester in Japan. On the 1425 RGB-D images in the test set (representing 51 unique potato tubers), our network achieved a completion accuracy of 2.8 mm on average. For volumetric estimation, the root mean squared error (RMSE) was 22.6 ml, and this was better than the RMSE of the linear regression (31.1 ml) and the base model (36.9 ml). We found that the RMSE can be further reduced to 18.2 ml when performing the 3D shape completion in the center of the RGB-D image. With an average 3D shape completion time of 10 milliseconds per tuber, we can conclude that CoRe++ is both fast and accurate enough to be implemented on an operational harvester for high-throughput potato yield estimation. CoRe++'s high-throughput and accurate processing allows it to be applied to other tuber, fruit and vegetable crops, thereby enabling versatile, accurate and real-time yield monitoring in precision agriculture. Our code, network weights and dataset are publicly available at https://github.com/UTokyo-FieldPhenomics-Lab/corepp.git.
Why it matches plant phenotyping methodsRGB-D画像からジャガイモ塊茎の3D形状を補完し、体積・収量を推定する手法を開発・検証しており、植物フェノタイピング手法が研究の中心である。
abstractwe developed a 3D shape completion network, called CoRe++, which can complete the 3D shape from RGB-D images.
Reproduction assets foundThe paper's abstract explicitly states that the authors' code, network weights, and the potato tuber RGB-D/3D point cloud dataset are publicly available at the authors' GitHub repository (UTokyo-FieldPhenomics-Lab/corepp), which is a paper-specific, public, actionable asset for the CoRe++ phenotyping analysis.Code · publicOur code, network weights and dataset are publicly available at https://github.com/UTokyo-FieldPhenomics-Lab/corepp.git .Open asset ↗UTokyo-FieldPhenomics-Lab/corepplines:1-93Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Creation of new annotated public datasets is crucial in helping advances in 3D computer vision and machine learning meet their full potential for automatic interpretation of 3D plant models. Despite the proliferation of deep neural network architectures for segmentation and phenotyping of 3D plant models in the last decade, the amount of data, and diversity in terms of species and data acquisition modalities are far from sufficient for evaluation of such tools for their generalization ability. To contribute to closing this gap, we introduce PLANesT-3D; a new annotated dataset of 3D color point clouds of plants. PLANesT-3D is composed of 34 point cloud models representing 34 real plants from three different plant species: \textit{Capsicum annuum}, \textit{Rosa kordana}, and \textit{Ribes rubrum}. Both semantic labels in terms of "leaf" and "stem", and organ instance labels were manually annotated for the full point clouds. PLANesT-3D introduces diversity to existing datasets by adding point clouds of two new species and providing 3D data acquired with the low-cost SfM/MVS technique as opposed to laser scanning or expensive setups. Point clouds reconstructed with SfM/MVS modality exhibit challenges such as missing data, variable density, and illumination variations. As an additional contribution, SP-LSCnet, a novel semantic segmentation method that is a combination of unsupervised superpoint extraction and a 3D point-based deep learning approach is introduced and evaluated on the new dataset. The advantages of SP-LSCnet over other deep learning methods are its modular structure and increased interpretability. Two existing deep neural network architectures, PointNet++ and RoseSegNet, were also tested on the point clouds of PLANesT-3D for semantic segmentation.
Why it matches plant phenotyping methods3D植物点群の注釈付きデータセットを構築し、植物器官のセマンティック・インスタンス分割手法を開発・評価しており、植物フェノタイピング手法が中心である。
abstractwe introduce PLANesT-3D; a new annotated dataset of 3D color point clouds of plants.
Reproduction assets foundThe paper introduces PLANesT-3D, an annotated 3D plant point cloud dataset, and SP-LSCnet segmentation code, both explicitly stated as publicly available at the authors' Aperta record and GitHub repository.Dataset · publicThe PLANesT-3D dataset is publicly available at https://aperta.ulakbim.gov.tr/record/286354 and https://github.com/visionlab-ogu/PLANesT-3D/tree/main/dataOpen asset ↗aperta.ulakbim.gov.tr · 286354lines:83-145Dataset · publicThe 2D color images for all the 34 plants together with their estimated camera poses and parameters are also open to the public to provide input data for recent 3D reconstruction techniques 3 3
3
The data is available at https://github.com/visionlab-ogu/PLANesT-3D/tree/main/data .Open asset ↗github.com/visionlab-ogu/PLANesT-3Dlines:494-505Code · publicThe code for SP-LSCnet is available at https://github.com/visionlab-ogu/PLANesT-3DOpen asset ↗github.com/visionlab-ogu/PLANesT-3Dlines:146-154Plant phenotyping relevance match · UnverifiedarXiv · checked 7 Sept 2026
BarleyPhysiological trait estimationGrowth / development / phenologyYield / yield components
Artificial Intelligence (AI) has emerged as a key driver of precision agriculture, facilitating enhanced crop productivity, optimized resource use, farm sustainability, and informed decision-making. Also, the expansion of genome sequencing technology has greatly increased crop genomic resources, deepening our understanding of genetic variation and enhancing desirable crop traits to optimize performance in various environments. There is increasing interest in using machine learning (ML) and deep learning (DL) algorithms for genotype-to-phenotype prediction due to their excellence in capturing complex interactions within large, high-dimensional datasets. In this work, we propose a new LSTM autoencoder-based model for barley genotype-to-phenotype prediction, specifically for flowering time and grain yield estimation, which could potentially help optimize yields and management practices. Our model outperformed the other baseline methods, demonstrating its potential in handling complex high-dimensional agricultural datasets and enhancing crop phenotype prediction performance.
Why it matches plant phenotyping methods大麦の開花期・穀粒収量という植物形質を予測するLSTMオートエンコーダモデルを新規開発し、ベースライン比較で性能検証しており、計算的な形質推定法が中心である。
abstractIn this work, we propose a new LSTM autoencoder-based model for barley genotype-to-phenotype prediction, specifically for flowering time and grain yield estimation
As the world population is expected to reach 10 billion by 2050, our agricultural production system needs to double its productivity despite a decline of human workforce in the agricultural sector. Autonomous robotic systems are one promising pathway to increase productivity by taking over labor-intensive manual tasks like fruit picking. To be effective, such systems need to monitor and interact with plants and fruits precisely, which is challenging due to the cluttered nature of agricultural environments causing, for example, strong occlusions. Thus, being able to estimate the complete 3D shapes of objects in presence of occlusions is crucial for automating operations such as fruit harvesting. In this paper, we propose the first publicly available 3D shape completion dataset for agricultural vision systems. We provide an RGB-D dataset for estimating the 3D shape of fruits. Specifically, our dataset contains RGB-D frames of single sweet peppers in lab conditions but also in a commercial greenhouse. For each fruit, we additionally collected high-precision point clouds that we use as ground truth. For acquiring the ground truth shape, we developed a measuring process that allows us to record data of real sweet pepper plants, both in the lab and in the greenhouse with high precision, and determine the shape of the sensed fruits. We release our dataset, consisting of almost 7,000 RGB-D frames belonging to more than 100 different fruits. We provide segmented RGB-D frames, with camera intrinsics to easily obtain colored point clouds, together with the corresponding high-precision, occlusion-free point clouds obtained with a high-precision laser scanner. We additionally enable evaluation of shape completion approaches on a hidden test set through a public challenge on a benchmark server.
Why it matches plant phenotyping methods果実の3D形状という植物器官の形態形質を対象に、RGB-D画像・高精度点群・評価用ベンチマークを構築しており、形状取得と推定の方法論が中心である。
abstractWe provide an RGB-D dataset for estimating the 3D shape of fruits.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Code · publicOur development toolkit including a data loader is available at:
https://github.com/PRBonn/shape_completion_toolkit for handling the dataset and computing metrics.Open asset ↗PRBonn/shape_completion_toolkitlines:55-81Code / dataset availability confirmedarXiv · checked 15 Sept 2026
High-throughput phenotyping refers to the non-destructive and efficient evaluation of plant phenotypes. In recent years, it has been coupled with machine learning in order to improve the process of phenotyping plants by increasing efficiency in handling large datasets and developing methods for the extraction of specific traits. Previous studies have developed methods to advance these challenges through the application of deep neural networks in tandem with automated cameras; however, the datasets being studied often excluded physical labels. In this study, we used a dataset provided by Oak Ridge National Laboratory with 1,672 images of Populus Trichocarpa with white labels displaying treatment (control or drought), block, row, position, and genotype. Optical character recognition (OCR) was used to read these labels on the plants, image segmentation techniques in conjunction with machine learning algorithms were used for morphological classifications, machine learning models were used to predict treatment based on those classifications, and analyzed encoded EXIF tags were used for the purpose of finding leaf size and correlations between phenotypes. We found that our OCR model had an accuracy of 94.31% for non-null text extractions, allowing for the information to be accurately placed in a spreadsheet. Our classification models identified leaf shape, color, and level of brown splotches with an average accuracy of 62.82%, and plant treatment with an accuracy of 60.08%. Finally, we identified a few crucial pieces of information absent from the EXIF tags that prevented the assessment of the leaf size. There was also missing information that prevented the assessment of correlations between phenotypes and conditions. However, future studies could improve upon this to allow for the assessment of these features.
Why it matches plant phenotyping methods植物画像からラベル情報を読み取り、画像分割・機械学習で葉形、色、斑点などの形態形質を抽出・分類する手法が研究の中心であり、植物フェノタイピング手法の開発・適用に該当する。
abstractimage segmentation techniques in conjunction with machine learning algorithms were used for morphological classifications
Reproduction assets foundThe paper's authors explicitly state that all analysis code (OCR label reading, leaf segmentation, morphology classification, treatment prediction) is publicly available under the MIT License on their GitHub repository. The underlying ORNL image dataset is not stated to be publicly available, so only the code asset is.Code · publicSince a pre-trained segmentation model (the SAM) was used in this study, researchers could attempt to build segmentation models fine-tuned to only recognize leaves, which could increase model efficiency and provide more consistent results.
6 Code Availability
All code is publicly available under the MIT License on GitHub here: https://github.com/vivaansinghvi07/smoky-mountain-data-comp .
Acknowledgements
We thank Dr. Ty Frazier at Oak Ridge National Laboratory for his helpful suggestions and mentoring throughout this project.
References
Arya et al. (2022)
Arya, S., Sandhu, K.S.,
Singh, J., Kumar, S.,
2022.
Deep learning: As the new frontier in high-throughput
plant phenotyping.
Euphytica 218Open asset ↗vivaansinghvi07/smoky-mountain-data-complines:272-401Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 13 Sept 2026
Field / plotChlorophyll fluorescenceRootWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionVisualization / data managementArchitecture / morphology / geometryRoot system architecture
Single-shot volumetric fluorescence (SVF) imaging offers a significant advantage over traditional imaging methods that require scanning across multiple axial planes as it can capture biological processes with high temporal resolution. The key challenges in SVF imaging include requiring sparsity constraints, eliminating depth ambiguity in the reconstruction, and maintaining high resolution across a large field of view. In this paper, we introduce the QuadraPol point spread function (PSF) combined with neural fields, a novel approach for SVF imaging. This method utilizes a custom polarizer at the back focal plane and a polarization camera to detect fluorescence, effectively encoding the 3D scene within a compact PSF without depth ambiguity. Additionally, we propose a reconstruction algorithm based on the neural fields technique that provides improved reconstruction quality compared to classical deconvolution methods. QuadraPol PSF, combined with neural fields, significantly reduces the acquisition time of a conventional fluorescence microscope by approximately 20 times and captures a 100 mm$^3$ cubic volume in one shot. We validate the effectiveness of both our hardware and algorithm through all-in-focus imaging of bacterial colonies on sand surfaces and visualization of plant root morphology. Our approach offers a powerful tool for advancing biological research and ecological studies.
Why it matches plant phenotyping methods植物根の形態を可視化する新規3D蛍光イメージング hardware と再構成アルゴリズムを開発・検証しており、植物フェノタイピング手法が中心である。
abstractIn this paper, we introduce the QuadraPol point spread function (PSF) combined with neural fields, a novel approach for SVF imaging.
Traditional field phenotyping methods are often manual, time-consuming, and destructive, posing a challenge for breeding progress. To address this bottleneck, robotics and automation technologies offer efficient sensing tools to monitor field evolution and crop development throughout the season. This study aimed to develop an autonomous ground robotic system for LiDAR-based field phenotyping in plant breeding trials. A Husky platform was equipped with a high-resolution three-dimensional (3D) laser scanner to collect in-field terrestrial laser scanning (TLS) data without human intervention. To automate the TLS process, a 3D ray casting analysis was implemented for optimal TLS site planning, and a route optimization algorithm was utilized to minimize travel distance during data collection. The platform was deployed in two cotton breeding fields for evaluation, where it autonomously collected TLS data. The system provided accurate pose information through RTK-GNSS positioning and sensor fusion techniques, with average errors of less than 0.6 cm for location and 0.38$^{\circ}$ for heading. The achieved localization accuracy allowed point cloud registration with mean point errors of approximately 2 cm, comparable to traditional TLS methods that rely on artificial targets and manual sensor deployment. This work presents an autonomous phenotyping platform that facilitates the quantitative assessment of plant traits under field conditions of both large agricultural fields and small breeding trials to contribute to the advancement of plant phenomics and breeding programs.
Why it matches plant phenotyping methods自律走行ロボットとLiDAR/TLSによる圃場フェノタイピング基盤を開発・評価しており、植物形質を定量評価するための取得・解析手法が研究の中心である。
abstractThis study aimed to develop an autonomous ground robotic system for LiDAR-based field phenotyping in plant breeding trials.
Field / plotNeRF / 3D Gaussian SplattingLiDAR / point cloudWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionArchitecture / morphology / geometry
We evaluate different Neural Radiance Fields (NeRFs) techniques for the 3D reconstruction of plants in varied environments, from indoor settings to outdoor fields. Traditional methods usually fail to capture the complex geometric details of plants, which is crucial for phenotyping and breeding studies. We evaluate the reconstruction fidelity of NeRFs in three scenarios with increasing complexity and compare the results with the point cloud obtained using LiDAR as ground truth. In the most realistic field scenario, the NeRF models achieve a 74.6% F1 score after 30 minutes of training on the GPU, highlighting the efficacy of NeRFs for 3D reconstruction in challenging environments. Additionally, we propose an early stopping technique for NeRF training that almost halves the training time while achieving only a reduction of 7.4% in the average F1 score. This optimization process significantly enhances the speed and efficiency of 3D reconstruction using NeRFs. Our findings demonstrate the potential of NeRFs in detailed and realistic 3D plant reconstruction and suggest practical approaches for enhancing the speed and efficiency of NeRFs in the 3D reconstruction process.
Why it matches plant phenotyping methods植物の3D形状再構成を対象にNeRF手法を評価し、LiDARを基準とした精度比較と早期停止による高速化を検証しているため、植物表現型取得法が中心である。
abstractWe evaluate different Neural Radiance Fields (NeRFs) techniques for the 3D reconstruction of plants in varied environments, from indoor settings to outdoor fields.
The process of estimating and counting tree density using only a single aerial or satellite image is a difficult task in the fields of photogrammetry and remote sensing. However, it plays a crucial role in the management of forests. The huge variety of trees in varied topography severely hinders tree counting models to perform well. The purpose of this paper is to propose a framework that is learnt from the source domain with sufficient labeled trees and is adapted to the target domain with only a limited number of labeled trees. Our method, termed as AdaTreeFormer, contains one shared encoder with a hierarchical feature extraction scheme to extract robust features from the source and target domains. It also consists of three subnets: two for extracting self-domain attention maps from source and target domains respectively and one for extracting cross-domain attention maps. For the latter, an attention-to-adapt mechanism is introduced to distill relevant information from different domains while generating tree density maps; a hierarchical cross-domain feature alignment scheme is proposed that progressively aligns the features from the source and target domains. We also adopt adversarial learning into the framework to further reduce the gap between source and target domains. Our AdaTreeFormer is evaluated on six designed domain adaptation tasks using three tree counting datasets, \ie Jiangsu, Yosemite, and London. Experimental results show that AdaTreeFormer significantly surpasses the state of the art, \eg in the cross domain from the Yosemite to Jiangsu dataset, it achieves a reduction of 15.9 points in terms of the absolute counting errors and an increase of 10.8\% in the accuracy of the detected trees' locations. The codes and datasets are available at https://github.com/HAAClassic/AdaTreeFormer.
Why it matches plant phenotyping methods単一の航空・衛星画像から樹木数・密度を推定する画像解析手法を開発し、複数データセットとドメイン適応タスクで評価しており、植物形質の取得が中心である。
abstractThe purpose of this paper is to propose a framework that is learnt from the source domain with sufficient labeled trees and is adapted to the target domain with only a limited number of labeled trees.
Reproduction assets foundThe paper publicly releases its AdaTreeFormer code and datasets, and evaluates on three publicly available tree-counting image/annotation datasets (Jiangsu, London, Yosemite) with explicit GitHub availability statements.Code · publicThe codes and datasets are available at https://github.com/HAAClassic/AdaTreeFormer .Open asset ↗HAAClassic/AdaTreeFormerlines:1-70Dataset · publicThis dataset encompasses 24 satellite images taken by the GaofenII satellite with a ground sample distance (GSD) of 0.8m (available at https://github.com/sddpltwanqiu/TreeCountNet/tree/main).Open asset ↗sddpltwanqiu/TreeCountNetlines:201-252Dataset · publicThis dataset consists of high-resolution images captured at 0.2m GSD from London, United Kingdom for training and testing (available at https://github.com/HAAClassic/TreeFormer/tree/main).Open asset ↗HAAClassic/TreeFormerlines:201-252Dataset · publicThe study area for this dataset revolves around Yosemite National Park, located in California, United States of America (available at https://github.com/nightonion/yosemite-tree-dataset ).Open asset ↗nightonion/yosemite-tree-datasetlines:201-252Plant phenotyping relevance match · UnverifiedarXiv · checked 7 Sept 2026
Crops for food, feed, fiber, and fuel are key natural resources for our society. Monitoring plants and measuring their traits is an important task in agriculture often referred to as plant phenotyping. Traditionally, this task is done manually, which is time- and labor-intensive. Robots can automate phenotyping providing reproducible and high-frequency measurements. Today's perception systems use deep learning to interpret these measurements, but require a substantial amount of annotated data to work well. Obtaining such labels is challenging as it often requires background knowledge on the side of the labelers. This paper addresses the problem of reducing the labeling effort required to perform leaf instance segmentation on 3D point clouds, which is a first step toward phenotyping in 3D. Separating all leaves allows us to count them and compute relevant traits as their areas, lengths, and widths. We propose a novel self-supervised task-specific pre-training approach to initialize the backbone of a network for leaf instance segmentation. We also introduce a novel automatic postprocessing that considers the difficulty of correctly segmenting the points close to the stem, where all the leaves petiole overlap. The experiments presented in this paper suggest that our approach boosts the performance over all the investigated scenarios. We also evaluate the embeddings to assess the quality of the fully unsupervised approach and see a higher performance of our domain-specific postprocessing.
Why it matches plant phenotyping methods3D点群から葉を個体別に分割し、面積・長さ・幅などの形質を抽出する計算手法を開発・評価しており、植物フェノタイピング手法が中心である。
abstractThis paper addresses the problem of reducing the labeling effort required to perform leaf instance segmentation on 3D point clouds, which is a first step toward phenotyping in 3D.
Agricultural robotics is an active research area due to global population growth and expectations of food and labor shortages. Robots can potentially help with tasks such as pruning, harvesting, phenotyping, and plant modeling. However, agricultural automation is hampered by the difficulty in creating high resolution 3D semantic maps in the field that would allow for safe manipulation and navigation. In this paper, we build toward solutions for this issue and showcase how the use of semantics and environmental priors can help in constructing accurate 3D maps for the target application of sorghum. Specifically, we 1) use sorghum seeds as semantic landmarks to build a visual Simultaneous Localization and Mapping (SLAM) system that enables us to map 78\\% of a sorghum range on average, compared to 38% with ORB-SLAM2; and 2) use seeds as semantic features to improve 3D reconstruction of a full sorghum panicle from images taken by a robotic in-hand camera.
Why it matches plant phenotyping methods植物の3D構造・器官形状を画像から再構成する手法開発が中心であり、単なるロボット位置推定に留まらず、ソルガム穂全体の3D再構成を扱っている。
abstractshowcase how the use of semantics and environmental priors can help in constructing accurate 3D maps for the target application of sorghum
Agricultural production is facing severe challenges in the next decades induced by climate change and the need for sustainability, reducing its impact on the environment. Advancements in field management through non-chemical weeding by robots in combination with monitoring of crops by autonomous unmanned aerial vehicles (UAVs) and breeding of novel and more resilient crop varieties are helpful to address these challenges. The analysis of plant traits, called phenotyping, is an essential activity in plant breeding, it however involves a great amount of manual labor. With this paper, we address the problem of automatic fine-grained organ-level geometric analysis needed for precision phenotyping. As the availability of real-world data in this domain is relatively scarce, we propose a novel dataset that was acquired using UAVs capturing high-resolution images of a real breeding trial containing 48 plant varieties and therefore covering great morphological and appearance diversity. This enables the development of approaches for autonomous phenotyping that generalize well to different varieties. Based on overlapping high-resolution images from multiple viewing angles, we compute photogrammetric dense point clouds and provide detailed and accurate point-wise labels for plants, leaves, and salient points as the tip and the base. Additionally, we include measurements of phenotypic traits performed by experts from the German Federal Plant Variety Office on the real plants, allowing the evaluation of new approaches not only on segmentation and keypoint detection but also directly on the downstream tasks. The provided labeled point clouds enable fine-grained plant analysis and support further progress in the development of automatic phenotyping approaches, but also enable further research in surface reconstruction, point cloud completion, and semantic interpretation of point clouds.
Why it matches plant phenotyping methods植物の器官レベル表現型解析を目的とするUAV画像由来の点群データセットで、植物・葉のラベルと専門家による形質測定を提供し、自動フェノタイピング手法の評価を可能にするため、方法論が中心である。
abstractwe propose a novel dataset that was acquired using UAVs capturing high-resolution images of a real breeding trial
Detecting and estimating size of apples during the early stages of growth is crucial for predicting yield, pest management, and making informed decisions related to crop-load management, harvest and post-harvest logistics, and marketing. Traditional fruit size measurement methods are laborious and timeconsuming. This study employs the state-of-the-art YOLOv8 object detection and instance segmentation algorithm in conjunction with geometric shape fitting techniques on 3D point cloud data to accurately determine the size of immature green apples (or fruitlet) in a commercial orchard environment. The methodology utilized two RGB-D sensors: Intel RealSense D435i and Microsoft Azure Kinect DK. Notably, the YOLOv8 instance segmentation models exhibited proficiency in immature green apple detection, with the YOLOv8m-seg model achieving the highest AP@0.5 and AP@0.75 scores of 0.94 and 0.91, respectively. Using the ellipsoid fitting technique on images from the Azure Kinect, we achieved an RMSE of 2.35 mm, MAE of 1.66 mm, MAPE of 6.15 mm, and an R-squared value of 0.9 in estimating the size of apple fruitlets. Challenges such as partial occlusion caused some error in accurately delineating and sizing green apples using the YOLOv8-based segmentation technique, particularly in fruit clusters. In a comparison with 102 outdoor samples, the size estimation technique performed better on the images acquired with Microsoft Azure Kinect than the same with Intel Realsense D435i. This superiority is evident from the metrics: the RMSE values (2.35 mm for Azure Kinect vs. 9.65 mm for Realsense D435i), MAE values (1.66 mm for Azure Kinect vs. 7.8 mm for Realsense D435i), and the R-squared values (0.9 for Azure Kinect vs. 0.77 for Realsense D435i).
Why it matches plant phenotyping methodsYOLOv8による検出・セグメンテーションと3D形状フィッティングを組み合わせ、リンゴ果実のサイズという植物器官形質を推定・検証する手法が研究の中心である。
abstractThis study employs the state-of-the-art YOLOv8 object detection and instance segmentation algorithm in conjunction with geometric shape fitting techniques on 3D point cloud data to accurately determine the size of immature green apples
Three-dimensional (3D) reconstruction of trees has always been a key task in precision forestry management and research. Due to the complex branch morphological structure of trees themselves and the occlusions from tree stems, branches and foliage, it is difficult to recreate a complete three-dimensional tree model from a two-dimensional image by conventional photogrammetric methods. In this study, based on tree images collected by various cameras in different ways, the Neural Radiance Fields (NeRF) method was used for individual tree reconstruction and the exported point cloud models are compared with point cloud derived from photogrammetric reconstruction and laser scanning methods. The results show that the NeRF method performs well in individual tree 3D reconstruction, as it has higher successful reconstruction rate, better reconstruction in the canopy area, it requires less amount of images as input. Compared with photogrammetric reconstruction method, NeRF has significant advantages in reconstruction efficiency and is adaptable to complex scenes, but the generated point cloud tends to be noisy and low resolution. The accuracy of tree structural parameters (tree height and diameter at breast height) extracted from the photogrammetric point cloud is still higher than those of derived from the NeRF point cloud. The results of this study illustrate the great potential of NeRF method for individual tree reconstruction, and it provides new ideas and research directions for 3D reconstruction and visualization of complex forest scenes.
Why it matches plant phenotyping methodsNeRFによる個体樹木の3D再構成を中心に、写真測量・レーザースキャンと比較検証し、樹高や胸高直径という植物形質を抽出しているため。
abstractthe Neural Radiance Fields (NeRF) method was used for individual tree reconstruction and the exported point cloud models are compared with point cloud derived from photogrammetric reconstruction and laser scanning methods
Automatic tree density estimation and counting using single aerial and satellite images is a challenging task in photogrammetry and remote sensing, yet has an important role in forest management. In this paper, we propose the first semisupervised transformer-based framework for tree counting which reduces the expensive tree annotations for remote sensing images. Our method, termed as TreeFormer, first develops a pyramid tree representation module based on transformer blocks to extract multi-scale features during the encoding stage. Contextual attention-based feature fusion and tree density regressor modules are further designed to utilize the robust features from the encoder to estimate tree density maps in the decoder. Moreover, we propose a pyramid learning strategy that includes local tree density consistency and local tree count ranking losses to utilize unlabeled images into the training process. Finally, the tree counter token is introduced to regulate the network by computing the global tree counts for both labeled and unlabeled images. Our model was evaluated on two benchmark tree counting datasets, Jiangsu, and Yosemite, as well as a new dataset, KCL-London, created by ourselves. Our TreeFormer outperforms the state of the art semi-supervised methods under the same setting and exceeds the fully-supervised methods using the same number of labeled images. The codes and datasets are available at https://github.com/HAAClassic/TreeFormer.
Why it matches plant phenotyping methods樹木の個体数・密度という植物状態を航空・衛星画像から推定する画像解析手法を開発し、複数データセットで評価しているため、植物フェノタイピング手法が中心である。
abstractAutomatic tree density estimation and counting using single aerial and satellite images is a challenging task
Reproduction assets foundThe paper's authors publicly release their analysis code and the KCL-London tree counting dataset via GitHub, and the paper's annotation workflow directly uses the public London Datastore local-authority-maintained trees dataset for tree locations.Code · publichmark tree counting datasets, Jiangsu, and Yosemite, as well as a new dataset, KCL-London, created by ourselves. Our TreeFormer outperforms the state of the art semi-supervised methods under the same setting and exceeds the fully-supervised methods using the same number of labeled images. The codes and datasets are available at https://github.com/HAAClassic/TreeFormer .
Index Terms:
Tree counting, semi-supervised model, transformer, pyramid learning strategy, remote sensing.
I Introduction
Trees are the pulse of the earth and are vital organisms in maintaining the ecological functioning and health of the planet [ 1 ] . Tree counting using high-resolution images is useful in various fields suOpen asset ↗HAAClassic/TreeFormerlines:1-71Dataset · publicmages are gathered and stitched together from Google Maps at 0.2 m ground sampling distance (GSD). The gathered images are divided into images with 1024 × 1024 pixels.
To aid the identification of tree locations and numbers of selected images, we employed the accessible tree locations of London in London Datastore website 1 1
1
https://data.london.gov.uk/dataset/local-authority-maintained-trees .
Although these data show the locations and species information for over 880,000 of London’s trees, the data mainly contains information on trees in the main streets and does not cover trees that are dense between houses or parks. We manually annotated the latter.
To this end, Global Mapper as geograOpen asset ↗lines:124-145Code / dataset availability confirmedarXiv · OpenAlex · checked 13 Sept 2026
Artificial intelligence applications enable farmers to optimize crop growth and production while reducing costs and environmental impact. Computer vision-based algorithms in particular, are commonly used for fruit segmentation, enabling in-depth analysis of the harvest quality and accurate yield estimation. In this paper, we propose TomatoDIFF, a novel diffusion-based model for semantic segmentation of on-plant tomatoes. When evaluated against other competitive methods, our model demonstrates state-of-the-art (SOTA) performance, even in challenging environments with highly occluded fruits. Additionally, we introduce Tomatopia, a new, large and challenging dataset of greenhouse tomatoes. The dataset comprises high-resolution RGB-D images and pixel-level annotations of the fruits.
Why it matches plant phenotyping methods植物上のトマト果実を画像からセグメンテーションする手法を開発・比較し、RGB-D画像と画素アノテーションのデータセットも提供しており、植物器官の状態・位置推定に関わる方法が中心である。
abstractwe propose TomatoDIFF, a novel diffusion-based model for semantic segmentation of on-plant tomatoes
Reproduction assets foundThe paper introduces TomatoDIFF and the Tomatopia dataset, with explicit public availability of source code and dataset at the authors' GitHub repository. It also trains/evaluates on the public Kaggle 'Tomato dataset' (andrewmvd/tomato-detection), which is a paper-specific public image dataset used directly in the phenCode · publicThe source code of TomatoDIFF and Tomatopia are available at https://github.com/MIvanovska/TomatoDIFF .Open asset ↗MIvanovska/TomatoDIFFlines:1-44Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 15 Sept 2026
Labor shortages in fruit crop production have prompted the development of mechanized and automated machines as alternatives to labor-intensive orchard operations such as harvesting, pruning, and thinning. Agricultural robots capable of identifying tree canopy parts and estimating geometric and topological parameters, such as branch diameter, length, and angles, can optimize crop yields through automated pruning and thinning platforms. In this study, we proposed a machine vision system to estimate canopy parameters in apple orchards and determine an optimal number of fruit for individual branches, providing a foundation for robotic pruning, flower thinning, and fruitlet thinning to achieve desired yield and quality.Using color and depth information from an RGB-D sensor (Microsoft Azure Kinect DK), a YOLOv8-based instance segmentation technique was developed to identify trunks and branches of apple trees during the dormant season. Principal Component Analysis was applied to estimate branch diameter (used to calculate limb cross-sectional area, or LCSA) and orientation. The estimated branch diameter was utilized to calculate LCSA, which served as an input for crop-load estimation, with larger LCSA values indicating a higher potential fruit-bearing capacity.RMSE for branch diameter estimation was 2.08 mm, and for crop-load estimation, 3.95. Based on commercial apple orchard management practices, the target crop-load (number of fruit) for each segmented branch was estimated with a mean absolute error (MAE) of 2.99 (ground truth crop-load was 6 apples per LCSA). This study demonstrated a promising workflow with high performance in identifying trunks and branches of apple trees in dynamic commercial orchard environments and integrating farm management practices into automated decision-making.
Why it matches plant phenotyping methodsRGB-D画像とYOLOv8を用いてリンゴ樹の枝形態(直径・方向)を抽出し、樹体の作物負荷を推定する手法を開発・評価しており、表現型取得が研究の中心である。
abstractIn this study, we proposed a machine vision system to estimate canopy parameters in apple orchards and determine an optimal number of fruit for individual branches
Multispectral / hyperspectralWhole plant / canopy / plot / field
The diversity of terrestrial vascular plants plays a key role in maintaining the stability and productivity of ecosystems. Airborne hyperspectral imaging has shown promise for measuring plant diversity remotely, but to operationalise these efforts over large regions we need to advance satellite-based alternatives. The advanced spectral and spatial specification of the recently launched DESIS (the DLR Earth Sensing Imaging Spectrometer) instrument provides a unique opportunity to test the potential for monitoring plant species diversity with spaceborne hyperspectral data. This study provides a quantitative assessment on the ability of DESIS hyperspectral data for predicting plant species richness in two different habitat types in southeast Australia. Spectral features were first extracted from the DESIS spectra, then regressed against on-ground estimates of plant species richness, with a two-fold cross validation scheme to assess the predictive performance. We tested and compared the effectiveness of Principal Component Analysis (PCA), Canonical Correlation Analysis (CCA), and Partial Least Squares analysis (PLS) for feature extraction, and Kernel Ridge Regression (KRR), Gaussian Process Regression (GPR), and Random Forest Regression (RFR) for species richness prediction. The best prediction results were $r=0.76$ and $\text{RMSE}=5.89$ for the Southern Tablelands region, and $r=0.68$ and $\text{RMSE}=5.95$ for the Snowy Mountains region. Relative importance analysis for the DESIS spectral bands showed that the red-edge, red, and blue spectral regions were more important for predicting plant species richness than the green bands and the near-infrared bands beyond red-edge. We also found that the DESIS hyperspectral data performed better than Sentinel-2 multispectral data in the prediction of plant species richness.
Why it matches plant phenotyping methodsDESISハイパースペクトルデータから植物種数を推定する特徴抽出・回帰手法を比較し、交差検証で性能評価しており、植物フェノタイピング手法が中心である。
abstractThis study provides a quantitative assessment on the ability of DESIS hyperspectral data for predicting plant species richness in two different habitat types in southeast Australia.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Dataset · publicFor on-ground measures of vascular plant species richness, we obtained plant community survey data from the NSW BioNet Vegetation Information System database [ Government, 2019 ] .Open asset ↗NSW BioNet Vegetation Information Systemlines:75-98Code / dataset availability confirmedarXiv · OpenAlex · checked 15 Sept 2026
In intensively managed forests in Europe, where forests are divided into stands of small size and may show heterogeneity within stands, a high spatial resolution (10 - 20 meters) is arguably needed to capture the differences in canopy height. In this work, we developed a deep learning model based on multi-stream remote sensing measurements to create a high-resolution canopy height map over the "Landes de Gascogne" forest in France, a large maritime pine plantation of 13,000 km$^2$ with flat terrain and intensive management. This area is characterized by even-aged and mono-specific stands, of a typical length of a few hundred meters, harvested every 35 to 50 years. Our deep learning U-Net model uses multi-band images from Sentinel-1 and Sentinel-2 with composite time averages as input to predict tree height derived from GEDI waveforms. The evaluation is performed with external validation data from forest inventory plots and a stereo 3D reconstruction model based on Skysat imagery available at specific locations. We trained seven different U-net models based on a combination of Sentinel-1 and Sentinel-2 bands to evaluate the importance of each instrument in the dominant height retrieval. The model outputs allow us to generate a 10 m resolution canopy height map of the whole "Landes de Gascogne" forest area for 2020 with a mean absolute error of 2.02 m on the Test dataset. The best predictions were obtained using all available satellite layers from Sentinel-1 and Sentinel-2 but using only one satellite source also provided good predictions. For all validation datasets in coniferous forests, our model showed better metrics than previous canopy height models available in the same region.
Why it matches plant phenotyping methodsSentinel/GEDI等のリモートセンシング画像から樹冠高を推定する深層学習手法を開発し、外部データで検証しているため、植物形質取得法が中心である。
abstractwe developed a deep learning model based on multi-stream remote sensing measurements to create a high-resolution canopy height map
Reproduction assets foundThe paper's primary phenotyping-relevant input is the GEDI L2A canopy height dataset (526,449 footprints over the Landes forest, 2020), explicitly downloaded from NASA's EarthDataSearch. This is a public, paper-specific sensor dataset directly used for the study's canopy height measurements and model training. No code,Dataset · publicwater bodies (Beck et al., 2020). Indeed, these
surfaces mirror the transmitted waveforms that have a pulse width of ~ 15 ns which
corresponds to a ~ 2.25 m wide waveform (Dubayah et al., 2020).
In total, 526,449 footprints from the GEDIv002 L2A product (Dubayah et al., 2021) were
downloaded from NASA’s EarthDataSearch website
(https://search.earthdata.nasa.gov/search) for this study, covering the entire area of interest
for 2020. Due to atmospheric perturbations, some waveforms could not be used to give
information on the vertical forest structure. Therefore, several filtering criteria were applied to
remove unusable waveforms: (1) When the quality_flag provided in the GEDI data was set toOpen asset ↗GEDIv002 L2Apdf-raw-page:6 lines:1-45Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 8 Sept 2026
Reliable and automated 3D plant shoot segmentation is a core prerequisite for the extraction of plant phenotypic traits at the organ level. Combining deep learning and point clouds can provide effective ways to address the challenge. However, fully supervised deep learning methods require datasets to be point-wise annotated, which is extremely expensive and time-consuming. In our work, we proposed a novel weakly supervised framework, Eff-3DPSeg, for 3D plant shoot segmentation. First, high-resolution point clouds of soybean were reconstructed using a low-cost photogrammetry system, and the Meshlab-based Plant Annotator was developed for plant point cloud annotation. Second, a weakly-supervised deep learning method was proposed for plant organ segmentation. The method contained: (1) Pretraining a self-supervised network using Viewpoint Bottleneck loss to learn meaningful intrinsic structure representation from the raw point clouds; (2) Fine-tuning the pre-trained model with about only 0.5% points being annotated to implement plant organ segmentation. After, three phenotypic traits (stem diameter, leaf width, and leaf length) were extracted. To test the generality of the proposed method, the public dataset Pheno4D was included in this study. Experimental results showed that the weakly-supervised network obtained similar segmentation performance compared with the fully-supervised setting. Our method achieved 95.1%, 96.6%, 95.8% and 92.2% in the Precision, Recall, F1-score, and mIoU for stem leaf segmentation and 53%, 62.8% and 70.3% in the AP, AP@25, and AP@50 for leaf instance segmentation. This study provides an effective way for characterizing 3D plant architecture, which will become useful for plant breeders to enhance selection processes.
Why it matches plant phenotyping methods3D点群による植物器官セグメンテーションと形質抽出手法を開発・検証しており、植物フェノタイピング手法が研究の中心である。
abstractReliable and automated 3D plant shoot segmentation is a core prerequisite for the extraction of plant phenotypic traits at the organ level.
In this paper, we present a method for creating high-quality 3D models of sorghum panicles for phenotyping in breeding experiments. This is achieved with a novel reconstruction approach that uses seeds as semantic landmarks in both 2D and 3D. To evaluate the performance, we develop a new metric for assessing the quality of reconstructed point clouds without having a ground-truth point cloud. Finally, a counting method is presented where the density of seed centers in the 3D model allows 2D counts from multiple views to be effectively combined into a whole-panicle count. We demonstrate that using this method to estimate seed count and weight for sorghum outperforms count extrapolation from 2D images, an approach used in most state of the art methods for seeds and grains of comparable size.
Why it matches plant phenotyping methodsソルガム穂の3D再構成と種子計数という植物形質取得手法を開発し、再構成品質評価指標と種子数・重量推定を検証しており、フェノタイピング手法が中心である。
abstractwe present a method for creating high-quality 3D models of sorghum panicles for phenotyping in breeding experiments
Reproduction assets foundThe paper's authors publicly release their sorghum panicle stereo-image dataset (camera poses, human-labeled seed segmentations, panicle weights, seed counts) via the CMU AIIRA resources page, which is an allowed URL.Dataset · publicection, some unremoved husks were counted as seeds by the counting machine despite manual efforts to separate seeds from husks. We expect the effect on the ground truth to be small. The stereo images, camera poses, human-labeled seed segmentations, panicle weights, and human-counted seed counts can be found in our dataset 3 3
3
https://labs.ri.cmu.edu/aiira/resources/ .
Figure 8: (a) 100 sorghum panicles from 10 different sorghum species. (b) Our data collection system, a stereo camera attached to the UR5 robot arm. (c) Seeds were manually stripped and (d) counted using a seed counting machine.
IV-B 3D Reconstruction Quality
We assess the effectiveness of our approach with ablation tests usiOpen asset ↗lines:141-165Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 13 Sept 2026
Active perception for fruit mapping and harvesting is a difficult task since occlusions occur frequently and the location as well as size of fruits change over time. State-of-the-art viewpoint planning approaches utilize computationally expensive ray casting operations to find good viewpoints aiming at maximizing information gain and covering the fruits in the scene. In this paper, we present a novel viewpoint planning approach that explicitly uses information about the predicted fruit shapes to compute targeted viewpoints that observe as yet unobserved parts of the fruits. Furthermore, we formulate the concept of viewpoint dissimilarity to reduce the sampling space for more efficient selection of useful, dissimilar viewpoints. Our simulation experiments with a UR5e arm equipped with an RGB-D sensor provide a quantitative demonstration of the efficacy of our iterative next best view planning method based on shape completion. In comparative experiments with a state-of-the-art viewpoint planner, we demonstrate improvement not only in the estimation of the fruit sizes, but also in their reconstruction, while significantly reducing the planning time. Finally, we show the viability of our approach for mapping sweet peppers plants with a real robotic system in a commercial glasshouse.
Why it matches plant phenotyping methods果実形状の再構成とサイズ推定を目的とする視点計画法を開発・比較評価しており、植物表現型取得が中心的な技術貢献である。
abstractwe present a novel viewpoint planning approach that explicitly uses information about the predicted fruit shapes to compute targeted viewpoints that observe as yet unobserved parts of the fruits.
We propose a novel hybrid cable-based robot with manipulator and camera for high-accuracy, medium-throughput plant monitoring in a vertical hydroponic farm and, as an example application, demonstrate non-destructive plant mass estimation. Plant monitoring with high temporal and spatial resolution is important to both farmers and researchers to detect anomalies and develop predictive models for plant growth. The availability of high-quality, off-the-shelf structure-from-motion (SfM) and photogrammetry packages has enabled a vibrant community of roboticists to apply computer vision for non-destructive plant monitoring. While existing approaches tend to focus on either high-throughput (e.g. satellite, unmanned aerial vehicle (UAV), vehicle-mounted, conveyor-belt imagery) or high-accuracy/robustness to occlusions (e.g. turn-table scanner or robot arm), we propose a middle-ground that achieves high accuracy with a medium-throughput, highly automated robot. Our design pairs the workspace scalability of a cable-driven parallel robot (CDPR) with the dexterity of a 4 degree-of-freedom (DoF) robot arm to autonomously image many plants from a variety of viewpoints. We describe our robot design and demonstrate it experimentally by collecting daily photographs of 54 plants from 64 viewpoints each. We show that our approach can produce scientifically useful measurements, operate fully autonomously after initial calibration, and produce better reconstructions and plant property estimates than those of over-canopy methods (e.g. UAV). As example applications, we show that our system can successfully estimate plant mass with a Mean Absolute Error (MAE) of 0.586g and, when used to perform hypothesis testing on the relationship between mass and age, produces p-values comparable to ground-truth data (p=0.0020 and p=0.0016, respectively).
Why it matches plant phenotyping methods植物の多視点画像取得、SfM再構成、質量推定を中核とするロボット型フェノタイピング手法の開発・実証であり、単なる生物学的測定ではない。
abstractWe describe our robot design and demonstrate it experimentally by collecting daily photographs of 54 plants from 64 viewpoints each.
In agriculture, the majority of vision systems perform still image classification. Yet, recent work has highlighted the potential of spatial and temporal cues as a rich source of information to improve the classification performance. In this paper, we propose novel approaches to explicitly capture both spatial and temporal information to improve the classification of deep convolutional neural networks. We leverage available RGB-D images and robot odometry to perform inter-frame feature map spatial registration. This information is then fused within recurrent deep learnt models, to improve their accuracy and robustness. We demonstrate that this can considerably improve the classification performance with our best performing spatial-temporal model (ST-Atte) achieving absolute performance improvements for intersection-over-union (IoU[%]) of 4.7 for crop-weed segmentation and 2.6 for fruit (sweet pepper) segmentation. Furthermore, we show that these approaches are robust to variable framerates and odometry errors, which are frequently observed in real-world applications.
Why it matches plant phenotyping methodsRGB-D画像とロボットオドメトリを用いて空間・時間情報を統合する画像解析手法を開発し、作物・果実のセグメンテーション性能を検証しているため、植物表現型取得手法が中心である。
abstractWe demonstrate that this can considerably improve the classification performance with our best performing spatial-temporal model (ST-Atte) achieving absolute performance improvements for intersection-over-union (IoU[%]) of 4.7 for crop-weed segmentation and 2.6 for fruit (sweet pepper) segmentation.
Robots in tomato greenhouses need to perceive the plant and plant parts accurately to automate monitoring, harvesting, and de-leafing tasks. Existing perception systems struggle with the high levels of occlusion in plants and often result in poor perception accuracy. One reason for this is because they use fixed cameras or predefined camera movements. Next-best-view (NBV) planning presents an alternate approach, in which the camera viewpoints are reasoned and strategically planned such that the perception accuracy is improved. However, existing NBV-planning algorithms are agnostic to the task-at-hand and give equal importance to all the plant parts. This strategy is inefficient for greenhouse tasks that require targeted perception of specific plant parts, such as the perception of leaf nodes for de-leafing. To improve targeted perception in complex greenhouse environments, NBV planning algorithms need an attention mechanism to focus on the task-relevant plant parts. In this paper, the role of attention in improving targeted perception using an attention-driven NBV planning strategy was investigated. Through simulation experiments using plants with high levels of occlusion and structural complexity, it was shown that focusing attention on task-relevant plant parts can significantly improve the speed and accuracy of 3D reconstruction. Further, with real-world experiments, it was shown that these benefits extend to complex greenhouse conditions with natural variation and occlusion, natural illumination, sensor noise, and uncertainty in camera poses. The results clearly indicate that using attention-driven NBV planning in greenhouses can significantly improve the efficiency of perception and enhance the performance of robotic systems in greenhouse crop production.
Why it matches plant phenotyping methods植物の3D再構成を効率化する注意機構付きNBV計画を開発・実験検証しており、植物部位の知覚・形状取得が中心的な方法論的貢献である。
abstractNext-best-view (NBV) planning presents an alternate approach, in which the camera viewpoints are reasoned and strategically planned such that the perception accuracy is improved.
Deep-learning-based image classification and object detection has been applied successfully to tree monitoring. However, studies of tree crowns and fallen trees, especially on flood inundated areas, remain largely unexplored. Detection of degraded tree trunks on natural environments such as water, mudflats, and natural vegetated areas is challenging due to the mixed colour image backgrounds. In this paper, Unmanned Aerial Vehicles (UAVs), or drones, with embedded RGB cameras were used to capture the fallen Acacia Xanthophloea trees from six designated plots around Lake Nakuru, Kenya. Motivated by the need to detect fallen trees around the lake, two well-established deep neural networks, i.e. Faster Region-based Convolution Neural Network (Faster R-CNN) and Retina-Net were used for fallen tree detection. A total of 7,590 annotations of three classes on 256 x 256 image patches were used for this study. Experimental results show the relevance of deep learning in this context, with Retina-Net model achieving 38.9% precision and 57.9% recall.
Why it matches plant phenotyping methodsUAV画像と深層学習によって個体レベルの倒木・劣化状態を検出する手法を中心に評価しており、植物状態の取得・推定方法が主要な貢献です。
abstractUnmanned Aerial Vehicles (UAVs), or drones, with embedded RGB cameras were used to capture the fallen Acacia Xanthophloea trees
Forest land plays a vital role in global climate, ecosystems, farming and human living environments. Therefore, forest biomass estimation methods are necessary to monitor changes in the forest structure and function, which are key data in natural resources research. Although accurate forest biomass measurements are important in forest inventory and assessments, high-density measurements that involve airborne light detection and ranging (LiDAR) at a low flight height in large mountainous areas are highly expensive. The objective of this study was to quantify the aboveground biomass (AGB) of a plateau mountainous forest reserve using a system that synergistically combines an unmanned aircraft system (UAS)-based digital aerial camera and LiDAR to leverage their complementary advantages. In this study, we utilized digital aerial photogrammetry (DAP), which has the unique advantages of speed, high spatial resolution, and low cost, to compensate for the deficiency of forestry inventory using UAS-based LiDAR that requires terrain-following flight for high-resolution data acquisition. Combined with the sparse LiDAR points acquired by using a high-altitude and high-speed UAS for terrain extraction, dense normalized DAP point clouds can be obtained to produce an accurate and high-resolution canopy height model (CHM). Based on the CHM and spectral attributes obtained from multispectral images, we estimated and mapped the AGB of the region of interest with considerable cost efficiency. Our study supports the development of predictive models for large-scale wall-to-wall AGB mapping by leveraging the complementarity between DAP and LiDAR measurements. This work also reveals the potential of utilizing a UAS-based digital camera and LiDAR synergistically in a plateau mountainous forest area.
Why it matches plant phenotyping methodsUASデジタルカメラとLiDARの融合により、森林キャノピー高と地上部バイオマスという植物群落形質を推定・マッピングする測定ワークフローが研究の中心である。
abstractThe objective of this study was to quantify the aboveground biomass (AGB) of a plateau mountainous forest reserve using a system that synergistically combines an unmanned aircraft system (UAS)-based digital aerial camera and LiDAR to leverage their complementary advantages.
Autonomous crop monitoring is a difficult task due to the complex structure of plants. Occlusions from leaves can make it impossible to obtain complete views about all fruits of, e.g., pepper plants. Therefore, accurately estimating the shape and volume of fruits from partial information is crucial to enable further advanced automation tasks such as yield estimation and automated fruit picking. In this paper, we present an approach for mapping fruits on plants and estimating their shape by matching superellipsoids. Our system segments fruits in images and uses their masks to generate point clouds of the fruits. To combine sequences of acquired point clouds, we utilize a real-time 3D mapping framework and build up a fruit map based on truncated signed distance fields. We cluster fruits from this map and use optimized superellipsoids for matching to obtain accurate shape estimates. In our experiments, we show in various simulated scenarios with a robotic arm equipped with an RGB-D camera that our approach can accurately estimate fruit volumes. Additionally, we provide qualitative results of estimated fruit shapes from data recorded in a commercial glasshouse environment.
Why it matches plant phenotyping methods果実の画像・RGB-Dデータから形状と体積を推定する手法が研究の中心であり、植物表現型の取得・抽出に直接該当する。
abstractestimating the shape and volume of fruits from partial information is crucial
Inspired by recent promising results in sim-to-real transfer in deep learning we built a realistic simulation environment combining a Robot Operating System (ROS)-compatible physics simulator (Gazebo) with Cycles, the realistic production rendering engine from Blender. The proposed simulator pipeline allows us to simulate near-realistic RGB-D images. To showcase the capabilities of the simulator pipeline we propose a case study that focuses on indoor robotic farming. We developed a solution for sweet pepper yield estimation task. Our approach to yield estimation starts with aerial robotics control and trajectory planning, combined with deep learning-based pepper detection, and a clustering approach for counting fruit. The results of this case study show that we can combine real time dynamic simulation with near realistic rendering capabilities to simulate complex robotic systems.
Why it matches plant phenotyping methods屋内農業向けにRGB-D画像シミュレーション基盤と、コショウ果実の検出・計数による収量推定手法を開発しており、植物形質取得・推定が中心的である。
abstractThe proposed simulator pipeline allows us to simulate near-realistic RGB-D images.
Obtaining 3D sensor data of complete plants or plant parts (e.g., the crop or fruit) is difficult due to their complex structure and a high degree of occlusion. However, especially for the estimation of the position and size of fruits, it is necessary to avoid occlusions as much as possible and acquire sensor information of the relevant parts. Global viewpoint planners exist that suggest a series of viewpoints to cover the regions of interest up to a certain degree, but they usually prioritize global coverage and do not emphasize the avoidance of local occlusions. On the other hand, there are approaches that aim at avoiding local occlusions, but they cannot be used in larger environments since they only reach a local maximum of coverage. In this paper, we therefore propose to combine a local, gradient-based method with global viewpoint planning to enable local occlusion avoidance while still being able to cover large areas. Our simulated experiments with a robotic arm equipped with a camera array as well as an RGB-D camera show that this combination leads to a significantly increased coverage of the regions of interest compared to just applying global coverage planning.
Why it matches plant phenotyping methods果実の位置・サイズ推定に必要な3Dセンサデータ取得を対象に、局所遮蔽回避と大域的視点計画を組み合わせる視点計画法を開発・評価しており、植物表現型取得が中心である。
abstractespecially for the estimation of the position and size of fruits, it is necessary to avoid occlusions as much as possible and acquire sensor information of the relevant parts
Reproduction assets foundThe paper's authors explicitly state that the source code of their combined local/global viewpoint planning system (used for fruit ROI coverage experiments) is publicly available on GitHub. OctoMap is a generic third-party library and is excluded.Code · publicThe source code of our system is available on GitHub 1 1
1
https://github.com/Eruvae/roi_viewpoint_planner .Open asset ↗Eruvae/roi_viewpoint_plannerlines:1-105Plant phenotyping relevance match · UnverifiedarXiv · checked 15 Sept 2026
We introduce a simple approach to understanding the relationship between single nucleotide polymorphisms (SNPs), or groups of related SNPs, and the phenotypes they control. The pipeline involves training deep convolutional neural networks (CNNs) to differentiate between images of plants with reference and alternate versions of various SNPs, and then using visualization approaches to highlight what the classification networks key on. We demonstrate the capacity of deep CNNs at performing this classification task, and show the utility of these visualizations on RGB imagery of biomass sorghum captured by the TERRA-REF gantry. We focus on several different genetic markers with known phenotypic expression, and discuss the possibilities of using this approach to uncover genotype x phenotype relationships.
Why it matches plant phenotyping methods植物画像からSNPに対応する表現型をCNNで分類・可視化する解析パイプラインが研究の中心であり、画像に基づく表現型抽出手法として適格です。
abstractThe pipeline involves training deep convolutional neural networks (CNNs) to differentiate between images of plants with reference and alternate versions of various SNPs, and then using visualization approaches to highlight what the classification networks key on.
Field / plotRGB / grayscaleMultispectral / hyperspectralThermalWhole plant / canopy / plot / field
A core objective of the TERRA-REF project was to generate an open-access reference dataset for the evaluation of sensing technologies to study plants under field conditions. The TERRA-REF program deployed a suite of high-resolution, cutting edge technology sensors on a gantry system with the aim of scanning 1 hectare (10$^4$) at around 1 mm$^2$ spatial resolution multiple times per week. The system contains co-located sensors including a stereo-pair RGB camera, a thermal imager, a laser scanner to capture 3D structure, and two hyperspectral cameras covering wavelengths of 300-2500nm. This sensor data is provided alongside over sixty types of traditional plant phenotype measurements that can be used to train new machine learning models. Associated weather and environmental measurements, information about agronomic management and experimental design, and the genomic sequences of hundreds of plant varieties have been collected and are available alongside the sensor and plant phenotype data. Over the course of four years and ten growing seasons, the TERRA-REF system generated over 1 PB of sensor data and almost 45 million files. The subset that has been released to the public domain accounts for two seasons and about half of the total data volume. This provides an unprecedented opportunity for investigations far beyond the core biological scope of the project. The focus of this paper is to provide the Computer Vision and Machine Learning communities an overview of the available data and some potential applications of this one of a kind data.
Why it matches plant phenotyping methods植物の高解像度マルチセンサーデータと植物表現型データを含む公開ベンチマーク/データセットを紹介し、コンピュータビジョンでの利用を主目的とするため、フェノタイピング手法・基盤として中心的です。
abstractgenerate an open-access reference dataset for the evaluation of sensing technologies to study plants under field conditions
Reproduction assets foundThe paper describes the TERRA-REF public domain release of plant phenotyping sensor data (RGB, thermal, laser scanner, hyperspectral, PSII) plus derived phenotypes, and explicitly points to public code repositories for the processing pipeline (terraref GitHub, PhytoOracle, AgPipeline) and a data access portal. All are,Dataset · publicprocessing, reviewing, curating, describing, and hosting the data.
Instead, we focused on an initial public release and plan to make new datasets available based on need.
Access to unpublished data can be requested from the authors, and as data are curated they will be added to subsequent versions of the public domain release ( https://terraref.org/data/access-data ).
In addition to hosting an archival copy of data on Dryad [ 16 ] , the
documentation includes instructions for browsing and accessing these
data through a variety of online portals. These portals provide access
to web user interfaces as well as databases, APIs, and R and Python
clients. In some cases it will be easier to acceOpen asset ↗lines:234-317Code · publicapproach described by Li et al . [ 18 ] .
Herritt et al . [ 14 , 13 ] demonstrate and provide software used in analysis of a sequence of images that capture plant fluorescence response to a pulse of light.
Most of the algorithms used to generate data products have not been published as papers but are made available on GitHub ( https://github.com/terraref ); code
used to release the data publication in 2020 is available on Zenodo [ 25 , 15 , 10 , 6 , 4 , 19 , 8 , 7 , 5 , 9 , 17 ] .
Pipeline development continues to support ongoing use of the field scanner as well as more general applications in plant sensing pipelines.
Recent advances have improved pipeline scalability and modulOpen asset ↗terrareflines:193-233Code · publiclant sensing pipelines.
Recent advances have improved pipeline scalability and modularity by adopting workflow tools and making use of heterogeneous computing environments.
The TERRA-REF computing pipeline has been adapted and extended for continuing use with the Field Scanner with the new name ”PhytoOracle” and is available at https://github.com/LyonsLab/PhytoOracle . Related work generalizing the pipeline for other phenomics applications has been released under the name ”AgPipeline” https://github.com/agpipeline with applications to aerial imaging described by Schnaufer et al . [ 22 ] .
All of these software are made available with permissive open source licenses on GitHub to enable accesOpen asset ↗PhytoOraclelines:193-233Code · publicnvironments.
The TERRA-REF computing pipeline has been adapted and extended for continuing use with the Field Scanner with the new name ”PhytoOracle” and is available at https://github.com/LyonsLab/PhytoOracle . Related work generalizing the pipeline for other phenomics applications has been released under the name ”AgPipeline” https://github.com/agpipeline with applications to aerial imaging described by Schnaufer et al . [ 22 ] .
All of these software are made available with permissive open source licenses on GitHub to enable access and community development.
Figure 4: Summary of public sensor datasets from Seasons 4 and 6. Each dot represents the dates for which a particular daOpen asset ↗agpipelinelines:193-233Plant phenotyping relevance match · UnverifiedarXiv · checked 15 Sept 2026
BlueberryStrawberryFruitClassificationGrowth / development / phenology
Automated technologies for quality inspection of fruits have attracted great interest in the food industry. The development of nondestructive mechanisms to assess the quality of individual fruit prior to sale may lead to an increase in overall product quality, value, and consequently, producer competitiveness. However, the existing methods have limitations. Herein, a texture sensor based on highly sensitive hair-like cilia receptors, to allow a quick quality evaluation of fruit is proposed. The texture sensor consists of up to 100 magnetized nanocomposite cilia attached to a chip with magnetoresistive sensors in a full Wheatstone bridge architecture. In this paper we demonstrate the use of ciliary sensors in scanning fruits (blueberries and strawberries) in different maturation stages. The contact of the cilia with the fruit skin provided qualitative information about its texture in terms of ripeness stage. Less mature fruits exhibited, on average, a highest peak voltage of 0.14 mV for blueberries and 0.12 mV for strawberries, while overripe fruits exhibited 0.58 mV and 0.56 mV, respectively. The results were confirmed by sensorial assessment of the fruit freshness, and therefore attesting the application potential of the sensing technology for fruit quality control.
Why it matches plant phenotyping methods果実の成熟度・テクスチャを直接推定する磁気式センサーを開発し、ブルーベリーとイチゴで検証しており、植物器官の状態取得が中心である。
abstractHerein, a texture sensor based on highly sensitive hair-like cilia receptors, to allow a quick quality evaluation of fruit is proposed.
Plant phenotyping, that is, the quantitative assessment of plant traits including growth, morphology, physiology, and yield, is a critical aspect towards efficient and effective crop management. Currently, plant phenotyping is a manually intensive and time consuming process, which involves human operators making measurements in the field, based on visual estimates or using hand-held devices. In this work, methods for automated grapevine phenotyping are developed, aiming to canopy volume estimation and bunch detection and counting. It is demonstrated that both measurements can be effectively performed in the field using a consumer-grade depth camera mounted onboard an agricultural vehicle.
Why it matches plant phenotyping methods消費者向け深度カメラを農業車両に搭載し、圃場でブドウの樹冠体積と房の検出・計数を自動化する手法を開発しており、表現型取得が研究の中心である。
abstractIn this work, methods for automated grapevine phenotyping are developed, aiming to canopy volume estimation and bunch detection and counting.
The cultivation of orchard meadows provides an ecological benefit for biodiversity, which is significantly higher than in intensively cultivated orchards. The goal of this research is to create a tree model to automatically determine possible pruning points for stand-alone trees within meadows. The algorithm which is presented here is capable of building a skeleton model based on a pre-segmented photogrammetric 3D point cloud. Good results were achieved in assigning the points to their leading branches and building a virtual tree model, reaching an overall accuracy of 95.19 %. This model provided the necessary information about the geometry of the tree for automated pruning.
Why it matches plant phenotyping methods3D点群から枝の骨格・樹体形状を推定する計算手法が中心で、単なる剪定対象の位置検出を超えて植物器官の構造形態を抽出している。
abstractThe algorithm which is presented here is capable of building a skeleton model based on a pre-segmented photogrammetric 3D point cloud.
Image-based yield detection in agriculture could raiseharvest efficiency and cultivation performance of farms. Following this goal, this research focuses on improving instance segmentation of field crops under varying environmental conditions. Five data sets of cabbage plants were recorded under varying lighting outdoor conditions. The images were acquired using a commercial mono camera. Additionally, depth information was generated out of the image stream with Structure-from-Motion (SfM). A Mask R-CNN was used to detect and segment the cabbage heads. The influence of depth information and different colour space representations were analysed. The results showed that depth combined with colour information leads to a segmentation accuracy increase of 7.1%. By describing colour information by colour spaces using light and saturation information combined with depth information, additional segmentation improvements of 16.5% could be reached. The CIELAB colour space combined with a depth information layer showed the best results achieving a mean average precision of 75.
Why it matches plant phenotyping methodsキャベツ頭部の画像セグメンテーションを対象に、深度情報と色空間の組合せを比較・改良し、精度を評価しているため、植物器官の取得・推定手法が中心である。
abstractthis research focuses on improving instance segmentation of field crops under varying environmental conditions
In this paper, we propose a novel deep learning method based on a Convolutional Neural Network (CNN) that simultaneously detects and geolocates plantation-rows while counting its plants considering highly-dense plantation configurations. The experimental setup was evaluated in a cornfield with different growth stages and in a Citrus orchard. Both datasets characterize different plant density scenarios, locations, types of crops, sensors, and dates. A two-branch architecture was implemented in our CNN method, where the information obtained within the plantation-row is updated into the plant detection branch and retro-feed to the row branch; which are then refined by a Multi-Stage Refinement method. In the corn plantation datasets (with both growth phases, young and mature), our approach returned a mean absolute error (MAE) of 6.224 plants per image patch, a mean relative error (MRE) of 0.1038, precision and recall values of 0.856, and 0.905, respectively, and an F-measure equal to 0.876. These results were superior to the results from other deep networks (HRNet, Faster R-CNN, and RetinaNet) evaluated with the same task and dataset. For the plantation-row detection, our approach returned precision, recall, and F-measure scores of 0.913, 0.941, and 0.925, respectively. To test the robustness of our model with a different type of agriculture, we performed the same task in the citrus orchard dataset. It returned an MAE equal to 1.409 citrus-trees per patch, MRE of 0.0615, precision of 0.922, recall of 0.911, and F-measure of 0.965. For citrus plantation-row detection, our approach resulted in precision, recall, and F-measure scores equal to 0.965, 0.970, and 0.964, respectively. The proposed method achieved state-of-the-art performance for counting and geolocating plants and plant-rows in UAV images from different types of crops.
Why it matches plant phenotyping methodsUAV画像から植物数を計数し、植栽列を検出・地理的位置特定するCNN手法の開発と比較評価が中心であり、植物形態・個体数の画像ベース表現型計測に該当する。
abstractwe propose a novel deep learning method based on a Convolutional Neural Network (CNN) that simultaneously detects and geolocates plantation-rows while counting its plants
Retrieval of vegetation properties from satellite and airborne optical data usually takes place after atmospheric correction, yet it is also possible to develop retrieval algorithms directly from top-of-atmosphere (TOA) radiance data. One of the key vegetation variables that can be retrieved from at-sensor TOA radiance data is the leaf area index (LAI) if algorithms account for variability in the atmosphere. We demonstrate the feasibility of LAI retrieval from Sentinel-2 (S2) TOA radiance data (L1C product) in a hybrid machine learning framework. To achieve this, the coupled leaf-canopy-atmosphere radiative transfer models PROSAIL-6S were used to simulate a look-up table (LUT) of TOA radiance data and associated input variables. This LUT was then used to train the Bayesian machine learning algorithms Gaussian processes regression (GPR) and variational heteroscedastic GPR (VHGPR). PROSAIL simulations were also used to train GPR and VHGPR models for LAI retrieval from S2 images at bottom-of-atmosphere (BOA) level (L2A product) for comparison purposes. The VHGPR models led to consistent LAI maps at BOA and TOA scale. We demonstrated that hybrid LAI retrieval algorithms can be developed from TOA radiance data given a cloud-free sky, thus without the need for atmospheric correction.
Why it matches plant phenotyping methodsSentinel-2のTOA放射輝度からLAIを推定する機械学習アルゴリズムを開発・比較しており、植物形質の取得手法が研究の中心である。
abstractWe demonstrate the feasibility of LAI retrieval from Sentinel-2 (S2) TOA radiance data (L1C product) in a hybrid machine learning framework.
Reproduction assets foundThe paper's hybrid LAI retrieval (GPR/VHGPR) was developed within the authors' ALG-ARTMO software framework, and code snippets/demos for GPR and VHGPR are publicly available from the authors' UV-ES soft regression page. Both are explicitly stated as freely downloadable in the supplied text. No paper-specific phenotype/Code · publicCode snippets and demos for both GPR, VHGPR and other machine learning regression algorithms is available from https://isp.uv.es/soft_regression.html .Open asset ↗isp.uv.es/soft_regression.htmllines:485-521Plant phenotyping relevance match · UnverifiedarXiv · checked 15 Sept 2026
This paper presents the algorithm developed in LSA-SAF (Satellite Application Facility for Land Surface Analysis) for the derivation of global vegetation parameters from the AVHRR (Advanced Very High-Resolution Radiometer) sensor onboard MetOp (Meteorological-Operational) satellites forming the EUMETSAT (European Organization for the Exploitation of Meteorological Satellites) Polar System (EPS). The suite of LSA-SAF EPS vegetation products includes the leaf area index (LAI), the fractional vegetation cover (FVC), and the fraction of absorbed photosynthetically active radiation (FAPAR). LAI, FAPAR, and FVC characterize the structure and the functioning of vegetation and are key parameters for a wide range of land-biosphere applications. The algorithm is based on a hybrid approach that blends the generalization capabilities offered by physical radiative transfer models with the accuracy and computational efficiency of machine learning methods. One major feature is the implementation of multi-output retrieval methods able to jointly and more consistently estimate all the biophysical parameters at the same time. We propose a multi-output Gaussian process regression (GPRmulti), which outperforms other considered methods over PROSAIL (coupling of PROSPECT and SAIL (Scattering by Arbitrary Inclined Leaves) radiative transfer models) EPS simulations. The global EPS products include uncertainty estimates taking into account the uncertainty captured by the retrieval method and input error propagation. The consistent generation and distribution of the EPS vegetation products will constitute a valuable tool for monitoring of earth surface dynamic processes.
Why it matches plant phenotyping methods衛星AVHRRから植物のLAI・FVC・FAPARを推定するアルゴリズムを開発し、機械学習手法の比較、マルチ出力推定、不確実性評価まで行っており、植物形質推定法が中心である。
abstractThis paper presents the algorithm developed in LSA-SAF (Satellite Application Facility for Land Surface Analysis) for the derivation of global vegetation parameters from the AVHRR (Advanced Very High-Resolution Radiometer) sensor onboard MetOp (Meteorological-Operational) satellites forming the EUMETSAT (European Organization for the Exploitation of Meteorological Satellites) Polar System (EPS).
Prediction of stress conditions is important for monitoring plant growth stages, disease detection, and assessment of crop yields. Multi-modal data, acquired from a variety of sensors, offers diverse perspectives and is expected to benefit the prediction process. We present several methods and strategies for abiotic stress prediction in banana plantlets, on a dataset acquired during a two and a half weeks period, of plantlets subject to four separate water and fertilizer treatments. The dataset consists of RGB and thermal images, taken once daily of each plant. Results are encouraging, in the sense that neural networks exhibit high prediction rates (over $90\%$ amongst four classes), in cases where there are hardly any noticeable features distinguishing the treatments, much higher than field experts can supply.
Why it matches plant phenotyping methodsRGB・熱画像からバナナ幼植物の非生物的ストレス状態を推定する画像解析手法とデータセットが研究の中心であり、植物状態のフェノタイピングに該当する。
abstractWe present several methods and strategies for abiotic stress prediction in banana plantlets
Modern agricultural applications require knowledge about the position and size of fruits on plants. However, occlusions from leaves typically make obtaining this information difficult. We present a novel viewpoint planning approach that builds up an octree of plants with labeled regions of interest (ROIs), i.e., fruits. Our method uses this octree to sample viewpoint candidates that increase the information around the fruit regions and evaluates them using a heuristic utility function that takes into account the expected information gain. Our system automatically switches between ROI targeted sampling and exploration sampling, which considers general frontier voxels, depending on the estimated utility. When the plants have been sufficiently covered with the RGB-D sensor, our system clusters the ROI voxels and estimates the position and size of the detected fruits. We evaluated our approach in simulated scenarios and compared the resulting fruit estimations with the ground truth. The results demonstrate that our combined approach outperforms a sampling method that does not explicitly consider the ROIs to generate viewpoints in terms of the number of discovered ROI cells. Furthermore, we show the real-world applicability by testing our framework on a robotic arm equipped with an RGB-D camera installed on an automated pipe-rail trolley in a capsicum glasshouse.
Why it matches plant phenotyping methods果実の位置・サイズという植物器官形質をRGB-Dセンサで取得・推定する視点計画法を開発し、シミュレーションと実環境で検証しており、フェノタイピング手法が中心である。
abstractWe present a novel viewpoint planning approach that builds up an octree of plants with labeled regions of interest (ROIs), i.e., fruits.
Reproduction assets foundThe paper's viewpoint-planning system source code and the simulated capsicum plant environments used in the experiments are publicly available on GitHub. OctoMap is a generic third-party library, not a paper-specific asset.Code · publicThe source code of our system is available on GitHub 1 1
1
https://github.com/Eruvae/roi_viewpoint_planner .Open asset ↗Eruvae/roi_viewpoint_plannerlines:73-113Plant phenotyping relevance match · UnverifiedarXiv · checked 9 Sept 2026
We present PATHoBot an autonomous crop surveying and intervention robot for glasshouse environments. The aim of this platform is to autonomously gather high quality data and also estimate key phenotypic parameters. To achieve this we retro-fit an off-the-shelf pipe-rail trolley with an array of multi-modal cameras, navigation sensors and a robotic arm for close surveying tasks and intervention. In this paper we describe PATHoBot design choices made to ensure proper operation in a commercial glasshouse environment. As a surveying platform we collect a number of datasets which include both sweet pepper and tomatoes. We show how PATHoBot enables novel surveillance approaches by first improving our previous work on fruit counting by incorporating wheel odometry and depth information. We find that by introducing re-projection and depth information we are able to achieve an absolute improvement of 20 points over the baseline technique in an "in the wild" situation. Finally, we present a 3D mapping case study, further showcasing PATHoBot's crop surveying capabilities.
Why it matches plant phenotyping methods温室作物の表現型パラメータ推定を目的とするロボット基盤を設計し、果実カウント手法を深度情報等で改善・評価しているため、表現型取得技術が中心である。
abstractThe aim of this platform is to autonomously gather high quality data and also estimate key phenotypic parameters.
Supervised learning is often used to count objects in images, but for counting small, densely located objects, the required image annotations are burdensome to collect. Counting plant organs for image-based plant phenotyping falls within this category. Object counting in plant images is further challenged by having plant image datasets with significant domain shift due to different experimental conditions, e.g. applying an annotated dataset of indoor plant images for use on outdoor images, or on a different plant species. In this paper, we propose a domain-adversarial learning approach for domain adaptation of density map estimation for the purposes of object counting. The approach does not assume perfectly aligned distributions between the source and target datasets, which makes it more broadly applicable within general object counting and plant organ counting tasks. Evaluation on two diverse object counting tasks (wheat spikelets, leaves) demonstrates consistent performance on the target datasets across different classes of domain shift: from indoor-to-outdoor images and from species-to-species adaptation.
Why it matches plant phenotyping methods植物器官数を画像から推定するドメイン適応・密度マップ推定法を開発し、コムギ小穂と葉の計数で評価しており、表現型取得手法が中心である。
abstractCounting plant organs for image-based plant phenotyping falls within this category.
Reproduction assets foundThe paper provides two paper-specific public assets: the authors' implementation code for the domain-adversarial counting model on GitHub, and the authors' newly created GWHD dot annotations deposited on figshare. Other URLs are cited prior-work datasets, not paper-specific assets.Code · publicAll experiments were performed on a GeForce RTX 2070 GPU with 8GB memory using the Pytorch framework. The implementation is available at: https://github.com/p2irc/UDA4POCOpen asset ↗p2irc/UDA4POClines:81-104Dataset · publicTo evaluate our method, we created dot annotations for 67 images from the GWHD which are used as ground truth. These annotations are made publicly available at https://doi.org/10.6084/m9.figshare.12652973.v2 .Open asset ↗10.6084/m9.figshare.12652973.v2lines:105-155Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Providing an accurate evaluation of palm tree plantation in a large region can bring meaningful impacts in both economic and ecological aspects. However, the enormous spatial scale and the variety of geological features across regions has made it a grand challenge with limited solutions based on manual human monitoring efforts. Although deep learning based algorithms have demonstrated potential in forming an automated approach in recent years, the labelling efforts needed for covering different features in different regions largely constrain its effectiveness in large-scale problems. In this paper, we propose a novel domain adaptive oil palm tree detection method, i.e., a Multi-level Attention Domain Adaptation Network (MADAN) to reap cross-regional oil palm tree counting and detection. MADAN consists of 4 procedures: First, we adopted a batch-instance normalization network (BIN) based feature extractor for improving the generalization ability of the model, integrating batch normalization and instance normalization. Second, we embedded a multi-level attention mechanism (MLA) into our architecture for enhancing the transferability, including a feature level attention and an entropy level attention. Then we designed a minimum entropy regularization (MER) to increase the confidence of the classifier predictions through assigning the entropy level attention value to the entropy penalty. Finally, we employed a sliding window-based prediction and an IOU based post-processing approach to attain the final detection results. We conducted comprehensive ablation experiments using three different satellite images of large-scale oil palm plantation area with six transfer tasks. MADAN improves the detection accuracy by 14.98% in terms of average F1-score compared with the Baseline method (without DA), and performs 3.55%-14.49% better than existing domain adaptation methods.
Why it matches plant phenotyping methods油ヤシ個体の計数・検出という植物形態/個体数形質を衛星画像から推定する手法を開発し、アブレーション実験と既存手法比較で検証しており、フェノタイピング手法が中心である。
abstractwe propose a novel domain adaptive oil palm tree detection method, i.e., a Multi-level Attention Domain Adaptation Network (MADAN) to reap cross-regional oil palm tree counting and detection.
Reproduction assets foundThe authors explicitly state that their code and datasets (satellite images and annotations used for oil palm tree detection) are publicly available on GitHub.Code · publiche oil palm tree detection performance across different
remotely sensed images acquired from different sensors, regions and dates, without using labeled samples in the
target region. Our MADAN is proposed for enhancing both the generalization capacity and the transferability of
our model. Our codes and datasets are available on https://github.com/rs-dl/MADAN. The major contributions of
our work are as follows:
(1) We propose an adaptive object detector named MADAN for oil palm tree counting and detection across different
satellite images, which is the first work for large-scale domain adaptive tree crown detection using multi-source and
multi-temporal remote sensing images.
(2) WeOpen asset ↗rs-dl/MADANpdf-raw-page:6 lines:1-22Plant phenotyping relevance match · UnverifiedarXiv · checked 9 Sept 2026
As a proposal-free approach, instance segmentation through pixel embedding learning and clustering is gaining more emphasis. Compared with bounding box refinement approaches, such as Mask R-CNN, it has potential advantages in handling complex shapes and dense objects. In this work, we propose a simple, yet highly effective, architecture for object-aware embedding learning. A distance regression module is incorporated into our architecture to generate seeds for fast clustering. At the same time, we show that the features learned by the distance regression module are able to promote the accuracy of learned object-aware embeddings significantly. By simply concatenating features of the distance regression module to the images as inputs of the embedding module, the mSBD scores on the CVPPP Leaf Segmentation Challenge can be further improved by more than 8% compared to the identical set-up without concatenation, yielding the best overall result amongst the leaderboard at CodaLab.
Why it matches plant phenotyping methodsCVPPP葉セグメンテーション課題を対象に、植物器官画像からのインスタンス分割を改善する埋め込み・距離回帰手法を開発しており、表現型抽出の計算手法が中心である。
titleImproving Pixel Embedding Learning through Intermediate Distance Regression Supervision for Instance Segmentation
In light of growing challenges in agriculture with ever growing food demand across the world, efficient crop management techniques are necessary to increase crop yield. Precision agriculture techniques allow the stakeholders to make effective and customized crop management decisions based on data gathered from monitoring crop environments. Plant phenotyping techniques play a major role in accurate crop monitoring. Advancements in deep learning have made previously difficult phenotyping tasks possible. This survey aims to introduce the reader to the state of the art research in deep plant phenotyping.
Why it matches plant phenotyping methods植物フェノタイピングにおけるコンピュータビジョンと深層学習を対象とするレビューであり、方法論の整理が中心です。
titleComputer Vision with Deep Learning for Plant Phenotyping in Agriculture: A Survey
Deep learning models have been successfully deployed for a diverse array of image-based plant phenotyping applications including disease detection and classification. However, successful deployment of supervised deep learning models requires large amount of labeled data, which is a significant challenge in plant science (and most biological) domains due to the inherent complexity. Specifically, data annotation is costly, laborious, time consuming and needs domain expertise for phenotyping tasks, especially for diseases. To overcome this challenge, active learning algorithms have been proposed that reduce the amount of labeling needed by deep learning models to achieve good predictive performance. Active learning methods adaptively select samples to annotate using an acquisition function to achieve maximum (classification) performance under a fixed labeling budget. We report the performance of four different active learning methods, (1) Deep Bayesian Active Learning (DBAL), (2) Entropy, (3) Least Confidence, and (4) Coreset, with conventional random sampling-based annotation for two different image-based classification datasets. The first image dataset consists of soybean [Glycine max L. (Merr.)] leaves belonging to eight different soybean stresses and a healthy class, and the second consists of nine different weed species from the field. For a fixed labeling budget, we observed that the classification performance of deep learning models with active learning-based acquisition strategies is better than random sampling-based acquisition for both datasets. The integration of active learning strategies for data annotation can help mitigate labelling challenges in the plant sciences applications particularly where deep domain knowledge is required.
Why it matches plant phenotyping methods植物画像フェノタイピングにおける能動学習手法を比較評価しており、ラベル付け削減と分類性能が中心的な方法論的貢献である。
titleHow useful is Active Learning for Image-based Plant Phenotyping?
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Code · publicAll the codes for the active learning approaches described in this work are available for the community at https://github.com/koushik-n/Active-Learning-Plant-Phenotyping.Open asset ↗koushik-n/Active-Learning-Plant-Phenotypinglines:108-129Plant phenotyping relevance match · UnverifiedarXiv · checked 9 Sept 2026
MaizeLeafCountingSegmentationTrackingGrowth / development / phenologyLeaf traits
Manual determination of plant phenotypic properties such as plant architecture, growth, and health is very time consuming and sometimes destructive. Automatic image analysis has become a popular approach. This research aims to identify the position (and number) of leaves from a temporal sequence of high-quality indoor images consisting of multiple views, focussing in particular of images of maize. The procedure used a segmentation on the images, using the convex hull to pick the best view at each time step, followed by a skeletonization of the corresponding image. To remove skeleton spurs, a discrete skeleton evolution pruning process was applied. Pre-existing statistics regarding maize development was incorporated to help differentiate between true leaves and false leaves. Furthermore, for each time step, leaves were matched to those of the previous and next three days using the graph-theoretic Hungarian algorithm. This matching algorithm can be used to both remove false positives, and also to predict true leaves, even if they were completely occluded from the image itself. The algorithm was evaluated using an open dataset consisting of 13 maize plants across 27 days from two different views. The total number of true leaves from the dataset was 1843, and our proposed techniques detect a total of 1690 leaves including 1674 true leaves, and only 16 false leaves, giving a recall of 90.8%, and a precision of 99.0%.
Why it matches plant phenotyping methodsトウモロコシ画像から葉の位置・数を自動抽出する画像処理と追跡アルゴリズムを開発し、公開データセットで精度評価しており、植物表現型取得手法が研究の中心である。
abstractThis research aims to identify the position (and number) of leaves from a temporal sequence of high-quality indoor images consisting of multiple views, focussing in particular of images of maize.
In this paper, we propose a deep learning framework for the automated counting and geolocation of palm trees from aerial images using convolutional neural networks. For this purpose, we collected aerial images in a palm tree Farm in the Kharj region, in Riyadh Saudi Arabia, using DJI drones, and we built a dataset of around 10,000 instances of palms trees. Then, we developed a convolutional neural network model using the state-of-the-art, Faster R-CNN algorithm. Furthermore, using the geotagged metadata of aerial images, we used photogrammetry concepts and distance corrections to detect the geographical location of detected palms trees automatically. This geolocation technique was tested on two different types of drones (DJI Mavic Pro, and Phantom 4 Pro), and was assessed to provide an average geolocation accuracy of 2.8m. This GPS tagging allows us to uniquely identify palm trees and count their number from a series of drone images, while correctly dealing with the issue of image overlapping. Moreover, it can be generalized to the geolocation of any other objects in UAV images.
Why it matches plant phenotyping methods航空画像からヤシ個体を自動検出・計数する画像解析手法を開発し、データセット構築と異なるドローンでの精度評価も行っており、植物個体数の取得が中心的な方法論的貢献である。
abstractwe propose a deep learning framework for the automated counting and geolocation of palm trees from aerial images using convolutional neural networks
The extraction of phenotypic traits is often very time and labour intensive. Especially the investigation in viticulture is restricted to an on-site analysis due to the perennial nature of grapevine. Traditionally skilled experts examine small samples and extrapolate the results to a whole plot. Thereby different grapevine varieties and training systems, e.g. vertical shoot positioning (VSP) and semi minimal pruned hedges (SMPH) pose different challenges. In this paper we present an objective framework based on automatic image analysis which works on two different training systems. The images are collected semi automatic by a camera system which is installed in a modified grape harvester. The system produces overlapping images from the sides of the plants. Our framework uses a convolutional neural network to detect single berries in images by performing a semantic segmentation. Each berry is then counted with a connected component algorithm. We compare our results with the Mask-RCNN, a state-of-the-art network for instance segmentation and with a regression approach for counting. The experiments presented in this paper show that we are able to detect green berries in images despite of different training systems. We achieve an accuracy for the berry detection of 94.0% in the VSP and 85.6% in the SMPH.
Why it matches plant phenotyping methodsブドウ果粒数という植物形質を、画像収集・セマンティックセグメンテーション・連結成分解析で自動抽出する方法が研究の中心であり、比較評価も実施している。
abstractIn this paper we present an objective framework based on automatic image analysis which works on two different training systems.
High-resolution cameras have become very helpful for plant phenotyping by providing a mechanism for tasks such as target versus background discrimination, and the measurement and analysis of fine-above-ground plant attributes. However, the acquisition of high-resolution (HR) imagery of plant roots is more challenging than above-ground data collection. Thus, an effective super-resolution (SR) algorithm is desired for overcoming resolution limitations of sensors, reducing storage space requirements, and boosting the performance of later analysis, such as automatic segmentation. We propose a SR framework for enhancing images of plant roots by using convolutional neural networks (CNNs). We compare three alternatives for training the SR model: i) training with non-plant-root images, ii) training with plant-root images, and iii) pretraining the model with non-plant-root images and fine-tuning with plant-root images. We demonstrate on a collection of publicly available datasets that the SR models outperform the basic bicubic interpolation even when trained with non-root datasets. Also, our segmentation experiments show that high performance on this task can be achieved independently of the SNR. Therefore, we conclude that the quality of the image enhancement depends on the application.
Why it matches plant phenotyping methods植物根画像の超解像化と後続の自動セグメンテーション性能を開発・比較する研究であり、根の表現型取得ワークフローにおける画像処理手法が中心である。
abstractWe propose a SR framework for enhancing images of plant roots by using convolutional neural networks (CNNs).
Deep neural networks have shown excellent performances in many real-world applications. Unfortunately, they may show "Clever Hans"-like behavior -- making use of confounding factors within datasets -- to achieve high performance. In this work, we introduce the novel learning setting of "explanatory interactive learning" (XIL) and illustrate its benefits on a plant phenotyping research task. XIL adds the scientist into the training loop such that she interactively revises the original model via providing feedback on its explanations. Our experimental results demonstrate that XIL can help avoiding Clever Hans moments in machine learning and encourages (or discourages, if appropriate) trust into the underlying model.
Why it matches plant phenotyping methods植物フェノタイピング課題を対象に、説明への研究者フィードバックを学習ループへ組み込む新しい機械学習手法を提案・実証しており、フェノタイピング解析手法が中心である。
abstractIn this work, we introduce the novel learning setting of "explanatory interactive learning" (XIL) and illustrate its benefits on a plant phenotyping research task.
Reproduction assets foundThe paper's plant phenotyping RGB/hyperspectral dataset is publicly deposited on TU Datalib, and the authors' analysis code (runnable Code Ocean capsule with pre-trained models reproducing figures/results) plus the user study materials are publicly available on GitHub. Generic benchmarks (Fashion-MNIST, PASCAL VOC) areCode · publiclable at
https://github.com/zalandoresearch/fashion-mnist . The PASCAL VOC2007 dataset is available at http://host.robots.ox.ac.uk/pascal/VOC/voc2007/ .
The RGB and hyperspectral data that support the findings of this study are available at https://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2278.4 and in the code repository https://codeocean.com/capsule/4559958/tree .
The user study is available at https://github.com/ml-research/xil/tree/master/Trust_Study .
Code availability
The code and a fully runnable capsule to reproduce the figures and results of this article, including pre-trained models, can be found at https://codeocean.com/capsule/4559958/tree .
Statement of ethical complianceOpen asset ↗codeocean · capsule/4559958lines:369-382Plant phenotyping relevance match · UnverifiedarXiv · checked 15 Sept 2026
We present techniques to measure crop heights using a 3D Light Detection and Ranging (LiDAR) sensor mounted on an Unmanned Aerial Vehicle (UAV). Knowing the height of plants is crucial to monitor their overall health and growth cycles, especially for high-throughput plant phenotyping. We present a methodology for extracting plant heights from 3D LiDAR point clouds, specifically focusing on plot-based phenotyping environments. We also present a toolchain that can be used to create phenotyping farms for use in Gazebo simulations. The tool creates a randomized farm with realistic 3D plant and terrain models. We conducted a series of simulations and hardware experiments in controlled and natural settings. Our algorithm was able to estimate the plant heights in a field with 112 plots with a root mean square error (RMSE) of 6.1 cm. This is the first such dataset for 3D LiDAR from an airborne robot over a wheat field. The developed simulation toolchain, algorithmic implementation, and datasets can be found on the GitHub repository located at https://github.com/hsd1121/PointCloudProcessing.
Why it matches plant phenotyping methodsUAV搭載3D LiDARによる作物高の抽出手法、シミュレーションツールチェーン、データセットを開発・評価しており、植物フェノタイピングが研究の中心である。
abstractWe present techniques to measure crop heights using a 3D Light Detection and Ranging (LiDAR) sensor mounted on an Unmanned Aerial Vehicle (UAV).
Panicle density of cereal crops such as wheat and sorghum is one of the main components for plant breeders and agronomists in understanding the yield of their crops. To phenotype the panicle density effectively, researchers agree there is a significant need for computer vision-based object detection techniques. Especially in recent times, research in deep learning-based object detection shows promising results in various agricultural studies. However, training such systems usually requires a lot of bounding-box labeled data. Since crops vary by both environmental and genetic conditions, acquisition of huge amount of labeled image datasets for each crop is expensive and time-consuming. Thus, to catalyze the widespread usage of automatic object detection for crop phenotyping, a cost-effective method to develop such automated systems is essential. We propose a point supervision based active learning approach for panicle detection in cereal crops. In our approach, the model constantly interacts with a human annotator by iteratively querying the labels for only the most informative images, as opposed to all images in a dataset. Our query method is specifically designed for cereal crops which usually tend to have panicles with low variance in appearance. Our method reduces labeling costs by intelligently leveraging low-cost weak labels (object centers) for picking the most informative images for which strong labels (bounding boxes) are required. We show promising results on two publicly available cereal crop datasets - Sorghum and Wheat. On Sorghum, 6 variants of our proposed method outperform the best baseline method with more than 55% savings in labeling time. Similarly, on Wheat, 3 variants of our proposed methods outperform the best baseline method with more than 50% of savings in labeling time.
Why it matches plant phenotyping methods穀粒穂の密度を対象とする画像ベースの検出・アクティブラーニング手法を開発し、複数作物データセットで性能とラベリングコストを評価しており、表現型取得手法が中心である。
abstractTo phenotype the panicle density effectively, researchers agree there is a significant need for computer vision-based object detection techniques.
Lodging, the permanent bending over of food crops, leads to poor plant growth and development. Consequently, lodging results in reduced crop quality, lowers crop yield, and makes harvesting difficult. Plant breeders routinely evaluate several thousand breeding lines, and therefore, automatic lodging detection and prediction is of great value aid in selection. In this paper, we propose a deep convolutional neural network (DCNN) architecture for lodging classification using five spectral channel orthomosaic images from canola and wheat breeding trials. Also, using transfer learning, we trained 10 lodging detection models using well-established deep convolutional neural network architectures. Our proposed model outperforms the state-of-the-art lodging detection methods in the literature that use only handcrafted features. In comparison to 10 DCNN lodging detection models, our proposed model achieves comparable results while having a substantially lower number of parameters. This makes the proposed model suitable for applications such as real-time classification using inexpensive hardware for high-throughput phenotyping pipelines. The GitHub repository at https://github.com/FarhadMaleki/LodgedNet contains code and models.
Why it matches plant phenotyping methodsUAV画像から作物の倒伏状態を推定するDCNN手法の開発・比較が中心であり、植物表現型の高スループット計測に直接関係する。
abstractwe propose a deep convolutional neural network (DCNN) architecture for lodging classification using five spectral channel orthomosaic images from canola and wheat breeding trials.
Phenotyping is the process of measuring an organism's observable traits. Manual phenotyping of crops is a labor-intensive, time-consuming, costly, and error prone process. Accurate, automated, high-throughput phenotyping can relieve a huge burden in the crop breeding pipeline. In this paper, we propose a scalable, high-throughput approach to automatically count and segment panicles (heads), a key phenotype, from aerial sorghum crop imagery. Our counting approach uses the image density map obtained from dot or region annotation as the target with a novel deep convolutional neural network architecture. We also propose a novel instance segmentation algorithm using the estimated density map, to identify the individual panicles in the presence of occlusion. With real Sorghum aerial images, we obtain a mean absolute error (MAE) of 1.06 for counting which is better than using well-known crowd counting approaches such as CCNN, MCNN and CSRNet models. The instance segmentation model also produces respectable results which will be ultimately useful in reducing the manual annotation workload for future data.
Why it matches plant phenotyping methodsソルガム穂(パンicles)の計数・個体セグメンテーションという植物形質を航空画像から自動抽出する手法の開発・評価が中心であり、明確なフェノタイピング方法論である。
abstractwe propose a scalable, high-throughput approach to automatically count and segment panicles (heads), a key phenotype, from aerial sorghum crop imagery.
Yield estimation and forecasting are of special interest in the field of grapevine breeding and viticulture. The number of harvested berries per plant is strongly correlated with the resulting quality. Therefore, early yield forecasting can enable a focused thinning of berries to ensure a high quality end product. Traditionally yield estimation is done by extrapolating from a small sample size and by utilizing historic data. Moreover, it needs to be carried out by skilled experts with much experience in this field. Berry detection in images offers a cheap, fast and non-invasive alternative to the otherwise time-consuming and subjective on-site analysis by experts. We apply fully convolutional neural networks on images acquired with the Phenoliner, a field phenotyping platform. We count single berries in images to avoid the error-prone detection of grapevine clusters. Clusters are often overlapping and can vary a lot in the size which makes the reliable detection of them difficult. We address especially the detection of white grapes directly in the vineyard. The detection of single berries is formulated as a classification task with three classes, namely 'berry', 'edge' and 'background'. A connected component algorithm is applied to determine the number of berries in one image. We compare the automatically counted number of berries with the manually detected berries in 60 images showing Riesling plants in vertical shoot positioned trellis (VSP) and semi minimal pruned hedges (SMPH). We are able to detect berries correctly within the VSP system with an accuracy of 94.0 \% and for the SMPH system with 85.6 \%.
Why it matches plant phenotyping methods画像と深層学習を用いてブドウ果粒数という収量関連形質を自動抽出し、手動計数と比較検証しており、植物フェノタイピング手法が中心です。
abstractBerry detection in images offers a cheap, fast and non-invasive alternative to the otherwise time-consuming and subjective on-site analysis by experts.
Sentinel-2 multi-spectral images collected over periods of several months were used to estimate vegetation height for Gabon and Switzerland. A deep convolutional neural network (CNN) was trained to extract suitable spectral and textural features from reflectance images and to regress per-pixel vegetation height. In Gabon, reference heights for training and validation were derived from airborne LiDAR measurements. In Switzerland, reference heights were taken from an existing canopy height model derived via photogrammetric surface reconstruction. The resulting maps have a mean absolute error (MAE) of 1.7 m in Switzerland and 4.3 m in Gabon (a root mean square error (RMSE) of 3.4 m and 5.6 m, respectively), and correctly estimate vegetation heights up to >50 m. They also show good qualitative agreement with existing vegetation height maps. Our work demonstrates that, given a moderate amount of reference data (i.e., 2000 km$^2$ in Gabon and $\approx$5800 km$^2$ in Switzerland), high-resolution vegetation height maps with 10 m ground sampling distance (GSD) can be derived at country scale from Sentinel-2 imagery.
Why it matches plant phenotyping methodsSentinel-2画像から植生高という明示的な植物形質をCNNで推定し、LiDAR等を用いて検証する測定手法の開発が中心であるため。
abstractA deep convolutional neural network (CNN) was trained to extract suitable spectral and textural features from reflectance images and to regress per-pixel vegetation height.
Images are used frequently in plant phenotyping to capture measurements. This chapter offers a repeatable method for capturing two-dimensional measurements of plant parts in field or laboratory settings using a variety of camera styles (cellular phone, DSLR), with the addition of a printed calibration pattern. The method is based on calibrating the camera using information available from the EXIF tags from the image, as well as visual information from the pattern. Code is provided to implement the method, as well as a dataset for testing. We include steps to verify protocol correctness by imaging an artifact. The use of this protocol for two-dimensional plant phenotyping will allow data capture from different cameras and environments, with comparison on the same physical scale. We abbreviate this method as CASS, for CAmera aS Scanner. Code and data is available at http://doi.org/10.5281/zenodo.3677473.
Why it matches plant phenotyping methods植物部位の2次元形質をカメラで測定する手法を開発し、校正・検証手順、コード、テストデータを提供しており、植物フェノタイピング手法が中心である。
abstractThis chapter offers a repeatable method for capturing two-dimensional measurements of plant parts in field or laboratory settings using a variety of camera styles (cellular phone, DSLR), with the addition of a printed calibration pattern.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Dataset · publicThe code and test datasets are provided in [ 16 ] . Within [ 16 ] , are some example data and two programs:
aruco-pattern-write and camera-as-scanner . To prepare for the experiments, download the example data and install the code (C++ code as well as a Docker image are provided).Open asset ↗lines:54-78Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 10 Sept 2026
The advent of plant phenomics, coupled with the wealth of genotypic data generated by next-generation sequencing technologies, provides exciting new resources for investigations into and improvement of complex traits. However, these new technologies also bring new challenges in quantitative genetics, namely, a need for the development of robust frameworks that can accommodate these high-dimensional data. In this chapter, we describe methods for the statistical analysis of high-throughput phenotyping (HTP) data with the goal of enhancing the prediction accuracy of genomic selection (GS). Following the Introduction in Section 1, Section 2 discusses field-based HTP, including the use of unmanned aerial vehicles and light detection and ranging, as well as how we can achieve increased genetic gain by utilizing image data derived from HTP. Section 3 considers extending commonly used GS models to integrate HTP data as covariates associated with the principal trait response, such as yield. Particular focus is placed on single-trait, multi-trait, and genotype by environment interaction models. One unique aspect of HTP data is that phenomics platforms often produce large-scale data with high spatial and temporal resolution for capturing dynamic growth, development, and stress responses. Section 4 discusses the utility of a random regression model for performing longitudinal GS. The chapter concludes with a discussion of some standing issues.
Why it matches plant phenotyping methodsHTPデータの統計解析と縦断的・多形質モデルを中心に扱う方法論的章であり、植物表現型データの解析ワークフローが主要内容である。
abstractIn this chapter, we describe methods for the statistical analysis of high-throughput phenotyping (HTP) data
Deep learning techniques involving image processing and data analysis are constantly evolving. Many domains adapt these techniques for object segmentation, instantiation and classification. Recently, agricultural industries adopted those techniques in order to bring automation to farmers around the globe. One analysis procedure required for automatic visual inspection in this domain is leaf count and segmentation. Collecting labeled data from field crops and greenhouses is a complicated task due to the large variety of crops, growth seasons, climate changes, phenotype diversity, and more, especially when specific learning tasks require a large amount of labeled data for training. Data augmentation for training deep neural networks is well established, examples include data synthesis, using generative semi-synthetic models, and applying various kinds of transformations. In this paper we propose a method that preserves the geometric structure of the data objects, thus keeping the physical appearance of the data-set as close as possible to imaged plants in real agricultural scenes. The proposed method provides state of the art results when applied to the standard benchmark in the field, namely, the ongoing Leaf Segmentation Challenge hosted by Computer Vision Problems in Plant Phenotyping.
Why it matches plant phenotyping methodsロゼット植物の葉のセグメンテーション・カウントという表現型抽出を対象に、幾何構造を保持するデータ拡張法を開発し、植物表現型ベンチマークで評価しているため、方法開発が中心です。
abstractIn this paper we propose a method that preserves the geometric structure of the data objects, thus keeping the physical appearance of the data-set as close as possible to imaged plants in real agricultural scenes.
Field / plotLiDAR / point cloudWhole plant / canopy / plot / fieldSegmentationYield / biomass estimationBiomass / plant weight
Developing a robust algorithm for automatic individual tree crown (ITC) detection from laser scanning datasets is important for tracking the responses of trees to anthropogenic change. Such approaches allow the size, growth and mortality of individual trees to be measured, enabling forest carbon stocks and dynamics to be tracked and understood. Many algorithms exist for structurally simple forests including coniferous forests and plantations. Finding a robust solution for structurally complex, species-rich tropical forests remains a challenge; existing segmentation algorithms often perform less well than simple area-based approaches when estimating plot-level biomass. Here we describe a Multi-Class Graph Cut (MCGC) approach to tree crown delineation. This uses local three-dimensional geometry and density information, alongside knowledge of crown allometries, to segment individual tree crowns from LiDAR point clouds. Our approach robustly identifies trees in the top and intermediate layers of the canopy, but cannot recognise small trees. From these three-dimensional crowns, we are able to measure individual tree biomass. Comparing these estimates to those from permanent inventory plots, our algorithm is able to produce robust estimates of hectare-scale carbon density, demonstrating the power of ITC approaches in monitoring forests. The flexibility of our method to add additional dimensions of information, such as spectral reflectance, make this approach an obvious avenue for future development and extension to other sources of three-dimensional data, such as structure from motion datasets.
Why it matches plant phenotyping methodsLiDAR点群から個体樹冠を抽出し、樹冠形状から個体バイオマスを推定する3次元フェノタイピング手法を開発・検証しており、方法が研究の中心である。
abstractHere we describe a Multi-Class Graph Cut (MCGC) approach to tree crown delineation.
A looming question that must be solved before robotic plant phenotyping capabilities can have significant impact to crop improvement programs is scalability. High Throughput Phenotyping (HTP) uses robotic technologies to analyze crops in order to determine species with favorable traits, however, the current practices rely on exhaustive coverage and data collection from the entire crop field being monitored under the breeding experiment. This works well in relatively small agricultural fields but can not be scaled to the larger ones, thus limiting the progress of genetics research. In this work, we propose an active learning algorithm to enable an autonomous system to collect the most informative samples in order to accurately learn the distribution of phenotypes in the field with the help of a Gaussian Process model. We demonstrate the superior performance of our proposed algorithm compared to the current practices on sorghum phenotype data collection.
Why it matches plant phenotyping methods植物表現型データ収集を効率化する能動学習・ガウス過程アルゴリズムを開発し、ソルガムの表現型データで既存手法と比較評価しており、表現型取得ワークフローが中心である。
abstractIn this work, we propose an active learning algorithm to enable an autonomous system to collect the most informative samples in order to accurately learn the distribution of phenotypes in the field with the help of a Gaussian Process model.
Reproduction assets foundThe authors explicitly open-sourced their code repository, simulation environment, and the sorghum phenotype dataset used in this paper, with a public GitHub URL.Code · publicWe have open-sourced our code repository, simulation environment and the sorghum dataset 1 1
1
Our github repository can be found at https://github.com/sumitsk/algp.git for the research community to carry out further work in this direction.Open asset ↗sumitsk/algp · sumitsk/algplines:55-75Dataset · publicOur simulation environment, code repository and the sorghum dataset are open-sourced and can be found at https://github.com/sumitsk/algp.git .Open asset ↗sumitsk/algp · sumitsk/algplines:180-200Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 13 Sept 2026
We tackle the challenging problem of creating full and accurate three dimensional reconstructions of botanical trees with the topological and geometric accuracy required for subsequent physical simulation, e.g. in response to wind forces. Although certain aspects of our approach would benefit from various improvements, our results exceed the state of the art especially in geometric and topological complexity and accuracy. Starting with two dimensional RGB image data acquired from cameras attached to drones, we create point clouds, textured triangle meshes, and a simulatable and skinned cylindrical articulated rigid body model. We discuss the pros and cons of each step of our pipeline, and in order to stimulate future research we make the raw and processed data from every step of the pipeline as well as the final geometric reconstructions publicly available.
Why it matches plant phenotyping methodsドローン画像から樹木の3次元形状・構造を再構成する手法とデータ公開が中心であり、植物の形態・樹冠構造を抽出する方法研究に該当する。
abstractcreating full and accurate three dimensional reconstructions of botanical trees with the topological and geometric accuracy required for subsequent physical simulation
We present a cheap, lightweight, and fast fruit counting pipeline that uses a single monocular camera. Our pipeline that relies only on a monocular camera, achieves counting performance comparable to state-of-the-art fruit counting system that utilizes an expensive sensor suite including LiDAR and GPS/INS on a mango dataset. Our monocular camera pipeline begins with a fruit detection component that uses a deep neural network. It then uses semantic structure from motion (SFM) to convert these detections into fruit counts by estimating landmark locations of the fruit in 3D, and using these landmarks to identify double counting scenarios. There are many benefits of developing a low cost and lightweight fruit counting system, including applicability to agriculture in developing countries, where monetary constraints or unstructured environments necessitate cheaper hardware solutions.
Why it matches plant phenotyping methods単眼カメラによる果実検出・3D位置推定・重複除去を統合し、果実数という植物器官の定量形質を抽出する手法が研究の中心であるため、植物フェノタイピング手法として採用する。
abstractWe present a cheap, lightweight, and fast fruit counting pipeline that uses a single monocular camera.
Measuring semantic traits for phenotyping is an essential but labor-intensive activity in horticulture. Researchers often rely on manual measurements which may not be accurate for tasks such as measuring tree volume. To improve the accuracy of such measurements and to automate the process, we consider the problem of building coherent three dimensional (3D) reconstructions of orchard rows. Even though 3D reconstructions of side views can be obtained using standard mapping techniques, merging the two side-views is difficult due to the lack of overlap between the two partial reconstructions. Our first main contribution in this paper is a novel method that utilizes global features and semantic information to obtain an initial solution aligning the two sides. Our mapping approach then refines the 3D model of the entire tree row by integrating semantic information common to both sides, and extracted using our novel robust detection and fitting algorithms. Next, we present a vision system to measure semantic traits from the optimized 3D model that is built from the RGB or RGB-D data captured by only a camera. Specifically, we show how canopy volume, trunk diameter, tree height and fruit count can be automatically obtained in real orchard environments. The experiment results from multiple datasets quantitatively demonstrate the high accuracy and robustness of our method.
Why it matches plant phenotyping methods果樹列の3D再構成と画像解析により、樹冠体積・幹径・樹高・果実数を自動測定する手法が研究の中心であり、植物表現型の取得方法を実環境で検証している。
abstractOur first main contribution in this paper is a novel method that utilizes global features and semantic information to obtain an initial solution aligning the two sides.
Automated segmentation of individual leaves of a plant in an image is a prerequisite to measure more complex phenotypic traits in high-throughput phenotyping. Applying state-of-the-art machine learning approaches to tackle leaf instance segmentation requires a large amount of manually annotated training data. Currently, the benchmark datasets for leaf segmentation contain only a few hundred labeled training images. In this paper, we propose a framework for leaf instance segmentation by augmenting real plant datasets with generated synthetic images of plants inspired by domain randomisation. We train a state-of-the-art deep learning segmentation architecture (Mask-RCNN) with a combination of real and synthetic images of Arabidopsis plants. Our proposed approach achieves 90% leaf segmentation score on the A1 test set outperforming the-state-of-the-art approaches for the CVPPP Leaf Segmentation Challenge (LSC). Our approach also achieves 81% mean performance over all five test datasets.
Why it matches plant phenotyping methods植物画像から個葉を自動セグメンテーションする手法を開発・評価しており、植物表現型抽出のための方法が中心的です。
abstractAutomated segmentation of individual leaves of a plant in an image is a prerequisite to measure more complex phenotypic traits in high-throughput phenotyping.
Reproduction assets foundThe paper's authors publicly released their generated synthetic Arabidopsis dataset (10,000 top-down images with 2D segmentation labels) used for training their leaf segmentation models, hosted on the CSIRO robotics databases page. The Matterport Mask_RCNN repository is a generic third-party library, and the CodaLab L5Dataset · publicOur generated synthetic dataset is publicly available at 3 3
3
https://research.csiro.au/robotics/databases . The synthetic dataset contains 10,000 top down images of synthetic Arabidopsis plants and their corresponding 2D segmentation labels.Open asset ↗lines:242-290Code / dataset availability confirmedOpenAlex · arXiv · checked 10 Sept 2026
We propose a new and, arguably, a very simple reduction of instance segmentation to semantic segmentation. This reduction allows to train feed-forward non-recurrent deep instance segmentation systems in an end-to-end fashion using architectures that have been proposed for semantic segmentation. Our approach proceeds by introducing a fixed number of labels (colors) and then dynamically assigning object instances to those labels during training (coloring). A standard semantic segmentation objective is then used to train a network that can color previously unseen images. At test time, individual object instances can be recovered from the output of the trained convolutional network using simple connected component analysis. In the experimental validation, the coloring approach is shown to be capable of solving diverse instance segmentation tasks arising in autonomous driving (the Cityscapes benchmark), plant phenotyping (the CVPPP leaf segmentation challenge), and high-throughput microscopy image analysis. The source code is publicly available: https://github.com/kulikovv/DeepColoring.
Why it matches plant phenotyping methodsインスタンスセグメンテーション手法の開発と実験検証が中心で、植物フェノタイピングの葉セグメンテーション課題に明示的に適用されている。
abstractWe propose a new and, arguably, a very simple reduction of instance segmentation to semantic segmentation.
Reproduction assets foundThe paper applies its Deep Coloring instance segmentation method to plant phenotyping (CVPPP leaf segmentation) and states its PyTorch implementation is publicly available on GitHub, enabling reproduction of the phenotyping analysis.Code · publicThe source code is publicly available: https://github.com/kulikovv/DeepColoring.Open asset ↗kulikovv/DeepColoringpdf-page:1 lines:1-64Plant phenotyping relevance match · UnverifiedOpenAlex · arXiv · checked 10 Sept 2026
Semantic labeling of 3D point clouds is important for the derivation of 3D models from real world scenarios in several economic fields such as building industry, facility management, town planning or heritage conservation. In contrast to these most common applications, we describe in this study the semantic labeling of 3D point clouds derived from plant organs by high-precision scanning. Our approach is optimized for the task of plant phenotyping with its very specific challenges and is employing a deep learning framework. Thereby, we report important experiences concerning detailed parameter initialization and optimization techniques. By evaluating our approach with challenging datasets we achieve state-of-the-art results without difficult and time consuming feature engineering as being necessary in traditional approaches to semantic labeling.
Why it matches plant phenotyping methods植物器官の3D点群を対象に、植物フェノタイピング向けのセマンティックラベリング手法を開発・評価しており、表現型取得・抽出手法が研究の中心である。
abstractwe describe in this study the semantic labeling of 3D point clouds derived from plant organs by high-precision scanning.
Measuring tree morphology for phenotyping is an essential but labor-intensive activity in horticulture. Researchers often rely on manual measurements which may not be accurate for example when measuring tree volume. Recent approaches on automating the measurement process rely on LIDAR measurements coupled with high-accuracy GPS. Usually each side of a row is reconstructed independently and then merged using GPS information. Such approaches have two disadvantages: (1) they rely on specialized and expensive equipment, and (2) since the reconstruction process does not simultaneously use information from both sides, side reconstructions may not be accurate. We also show that standard loop closure methods do not necessarily align tree trunks well. In this paper, we present a novel vision system that employs only an RGB-D camera to estimate morphological parameters. A semantics-based mapping algorithm merges the two-sides 3D models of tree rows, where integrated semantic information is obtained and refined by robust fitting algorithms. We focus on measuring tree height, canopy volume and trunk diameter from the optimized 3D model. Experiments conducted in real orchards quantitatively demonstrate the accuracy of our method.
Why it matches plant phenotyping methodsRGB-Dカメラと意味ベースの3Dマッピングにより樹体形態を推定する手法を開発し、果樹園で精度検証しているため、植物フェノタイピング手法が中心である。
abstractIn this paper, we present a novel vision system that employs only an RGB-D camera to estimate morphological parameters.
We present a novel fruit counting pipeline that combines deep segmentation, frame to frame tracking, and 3D localization to accurately count visible fruits across a sequence of images. Our pipeline works on image streams from a monocular camera, both in natural light, as well as with controlled illumination at night. We first train a Fully Convolutional Network (FCN) and segment video frame images into fruit and non-fruit pixels. We then track fruits across frames using the Hungarian Algorithm where the objective cost is determined from a Kalman Filter corrected Kanade-Lucas-Tomasi (KLT) Tracker. In order to correct the estimated count from tracking process, we combine tracking results with a Structure from Motion (SfM) algorithm to calculate relative 3D locations and size estimates to reject outliers and double counted fruit tracks. We evaluate our algorithm by comparing with ground-truth human-annotated visual counts. Our results demonstrate that our pipeline is able to accurately and reliably count fruits across image sequences, and the correction step can significantly improve the counting accuracy and robustness. Although discussed in the context of fruit counting, our work can extend to detection, tracking, and counting of a variety of other stationary features of interest such as leaf-spots, wilt, and blossom.
Why it matches plant phenotyping methods果実数を画像から抽出する深層学習・追跡・SfM統合パイプラインを開発し、アノテーション済み計数で精度検証しており、植物表現型取得が中心である。
abstractWe present a novel fruit counting pipeline that combines deep segmentation, frame to frame tracking, and 3D localization to accurately count visible fruits across a sequence of images.
The berry size is one of the most important fruit traits in grapevine breeding. Non-invasive, image-based phenotyping promises a fast and precise method for the monitoring of the grapevine berry size. In the present study an automated image analyzing framework was developed in order to estimate the size of grapevine berries from images in a high-throughput manner. The framework includes (i) the detection of circular structures which are potentially berries and (ii) the classification of these into the class 'berry' or 'non-berry' by utilizing a conditional random field. The approach used the concept of a one-class classification, since only the target class 'berry' is of interest and needs to be modeled. Moreover, the classification was carried out by using an automated active learning approach, i.e no user interaction is required during the classification process and in addition, the process adapts automatically to changing image conditions, e.g. illumination or berry color. The framework was tested on three datasets consisting in total of 139 images. The images were taken in an experimental vineyard at different stages of grapevine growth according to the BBCH scale. The mean berry size of a plant estimated by the framework correlates with the manually measured berry size by $0.88$.
Why it matches plant phenotyping methodsブドウ果実サイズという植物形質を画像から高スループット推定する解析フレームワークを開発し、手測定との相関で検証しており、フェノタイピング手法が研究の中心である。
abstractan automated image analyzing framework was developed in order to estimate the size of grapevine berries from images in a high-throughput manner.
Quantification of physiological changes in plants can capture different drought mechanisms and assist in selection of tolerant varieties in a high throughput manner. In this context, an accurate 3D model of plant canopy provides a reliable representation for drought stress characterization in contrast to using 2D images. In this paper, we propose a novel end-to-end pipeline including 3D reconstruction, segmentation and feature extraction, leveraging deep neural networks at various stages, for drought stress study. To overcome the high degree of self-similarities and self-occlusions in plant canopy, prior knowledge of leaf shape based on features from deep siamese network are used to construct an accurate 3D model using structure from motion on wheat plants. The drought stress is characterized with a deep network based feature aggregation. We compare the proposed methodology on several descriptors, and show that the network outperforms conventional methods.
Why it matches plant phenotyping methods植物キャノピーの3D再構成、セグメンテーション、特徴抽出、乾燥ストレス評価を統合した画像ベースの表現型解析パイプラインが研究の中心であるため。
abstractwe propose a novel end-to-end pipeline including 3D reconstruction, segmentation and feature extraction
In recent years, there has been an increasing interest in image-based plant phenotyping, applying state-of-the-art machine learning approaches to tackle challenging problems, such as leaf segmentation (a multi-instance problem) and counting. Most of these algorithms need labelled data to learn a model for the task at hand. Despite the recent release of a few plant phenotyping datasets, large annotated plant image datasets for the purpose of training deep learning algorithms are lacking. One common approach to alleviate the lack of training data is dataset augmentation. Herein, we propose an alternative solution to dataset augmentation for plant phenotyping, creating artificial images of plants using generative neural networks. We propose the Arabidopsis Rosette Image Generator (through) Adversarial Network: a deep convolutional network that is able to generate synthetic rosette-shaped plants, inspired by DCGAN (a recent adversarial network model using convolutional layers). Specifically, we trained the network using A1, A2, and A4 of the CVPPP 2017 LCC dataset, containing Arabidopsis Thaliana plants. We show that our model is able to generate realistic 128x128 colour images of plants. We train our network conditioning on leaf count, such that it is possible to generate plants with a given number of leaves suitable, among others, for training regression based models. We propose a new Ax dataset of artificial plants images, obtained by our ARIGAN. We evaluate this new dataset using a state-of-the-art leaf counting algorithm, showing that the testing error is reduced when Ax is used as part of the training data.
Why it matches plant phenotyping methods植物表現型解析用の合成画像生成ネットワークを開発し、葉数を条件付けたデータセットを作成・評価しており、表現型取得・解析ワークフローの技術的貢献が中心である。
abstractWe propose a new Ax dataset of artificial plants images, obtained by our ARIGAN.
Reproduction assets foundThe paper's authors publicly released the Ax dataset of 57 synthetic Arabidopsis plant images generated by ARIGAN (with leaf-count annotations in a CSV), which directly reproduces the paper's phenotyping data contribution. The CVPPP 2017 LCC dataset is the training input but is cited prior work, not a paper-specific.Dataset · publicr quantitative experiments show that the extension of the training dataset with the images in Ax improved the testing error and reduced overfitting. We run a 4-fold cross validation experiment on A4 dataset. Evaluation metrics of our experiments are reported in Table 1 . Our synthetic dataset Ax is available to download at \url http://www.valeriogiuffrida.academy/ax.
Acknowledgements
This work was supported by The Alan Turing Institute under the EPSRC grant EP/N510129/1, and also by the BBSRC grant BB/P023487/1.
References
[1]
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. J. Goodfellow, A. Bergeron,
N. Bouchard, and Y. Bengio.
Theano: new features and speed improvements.
Deep LearninOpen asset ↗Axlines:101-153Code / dataset availability confirmedarXiv · checked 10 Sept 2026
In this paper, we investigate the problem of counting rosette leaves from an RGB image, an important task in plant phenotyping. We propose a data-driven approach for this task generalized over different plant species and imaging setups. To accomplish this task, we use state-of-the-art deep learning architectures: a deconvolutional network for initial segmentation and a convolutional network for leaf counting. Evaluation is performed on the leaf counting challenge dataset at CVPPP-2017. Despite the small number of training samples in this dataset, as compared to typical deep learning image sets, we obtain satisfactory performance on segmenting leaves from the background as a whole and counting the number of leaves using simple data augmentation strategies. Comparative analysis is provided against methods evaluated on the previous competition datasets. Our framework achieves mean and standard deviation of absolute count difference of 1.62 and 2.30 averaged over all five test datasets.
Why it matches plant phenotyping methodsロゼット葉の画像から葉数を推定する深層学習手法を提案し、セグメンテーションと葉数カウントをデータセットで評価しており、植物表現型取得手法が中心である。
abstractcounting rosette leaves from an RGB image, an important task in plant phenotyping
Reproduction assets foundThe paper's authors explicitly state their leaf counting/segmentation code is publicly available on GitHub, and the CVPPP2017 Leaf Counting Challenge dataset used for all experiments is publicly hosted at plant-phenotyping.org.Code · publicCode is publicly available here. 1 1Open asset ↗lines:215-305Code / dataset availability confirmedarXiv · checked 14 Sept 2026
Phenomics is an emerging branch of modern biology that uses high throughput phenotyping tools to capture multiple environmental and phenotypic traits, often at massive spatial and temporal scales. The resulting high dimensional data represent a treasure trove of information for providing an in-depth understanding of how multiple factors interact and contribute to the overall growth and behavior of different genotypes. However, computational tools that can parse through such complex data and aid in extracting plausible hypotheses are currently lacking. In this paper, we present Hyppo-X, a new algorithmic approach to visually explore complex phenomics data and in the process characterize the role of environment on phenotypic traits. We model the problem as one of unsupervised structure discovery, and use emerging principles from algebraic topology and graph theory for discovering higher-order structures of complex phenomics data. We present an open source software which has interactive visualization capabilities to facilitate data navigation and hypothesis formulation. We test and evaluate Hyppo-X on two real-world plant (maize) data sets. Our results demonstrate the ability of our approach to delineate divergent subpopulation-level behavior. Notably, our approach shows how environmental factors could influence phenotypic behavior, and how that effect varies across different genotypes and different time scales. To the best of our knowledge, this effort provides one of the first approaches to systematically formalize the problem of hypothesis extraction for phenomics data. Considering the infancy of the phenomics field, tools that help users explore complex data and extract plausible hypotheses in a data-guided manner will be critical to future advancements in the use of such data.
Why it matches plant phenotyping methods植物フェノミクスデータを探索・解析するアルゴリズムとオープンソースソフトウェアを開発し、植物データセットで評価しているため、フェノタイピング解析手法が中心である。
abstractIn this paper, we present Hyppo-X, a new algorithmic approach to visually explore complex phenomics data
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Code · publicThe tool is available as open source in the GitHub repository [ 18 ] .Open asset ↗lines:185-261Plant phenotyping relevance match · UnverifiedarXiv · OpenAlex · checked 10 Sept 2026
Accurately counting maize tassels is important for monitoring the growth status of maize plants. This tedious task, however, is still mainly done by manual efforts. In the context of modern plant phenotyping, automating this task is required to meet the need of large-scale analysis of genotype and phenotype. In recent years, computer vision technologies have experienced a significant breakthrough due to the emergence of large-scale datasets and increased computational resources. Naturally image-based approaches have also received much attention in plant-related studies. Yet a fact is that most image-based systems for plant phenotyping are deployed under controlled laboratory environment. When transferring the application scenario to unconstrained in-field conditions, intrinsic and extrinsic variations in the wild pose great challenges for accurate counting of maize tassels, which goes beyond the ability of conventional image processing techniques. This calls for further robust computer vision approaches to address in-field variations. This paper studies the in-field counting problem of maize tassels. To our knowledge, this is the first time that a plant-related counting problem is considered using computer vision technologies under unconstrained field-based environment.
Why it matches plant phenotyping methodsトウモロコシ雄穂数という植物形質を、野外画像から自動推定するコンピュータビジョン手法の開発が中心である。
abstractThis paper studies the in-field counting problem of maize tassels.
Machine vision for plant phenotyping is an emerging research area for producing high throughput in agriculture and crop science applications. Since 2D based approaches have their inherent limitations, 3D plant analysis is becoming state of the art for current phenotyping technologies. We present an automated system for analyzing plant growth in indoor conditions. A gantry robot system is used to perform scanning tasks in an automated manner throughout the lifetime of the plant. A 3D laser scanner mounted as the robot's payload captures the surface point cloud data of the plant from multiple views. The plant is monitored from the vegetative to reproductive stages in light/dark cycles inside a controllable growth chamber. An efficient 3D reconstruction algorithm is used, by which multiple scans are aligned together to obtain a 3D mesh of the plant, followed by surface area and volume computations. The whole system, including the programmable growth chamber, robot, scanner, data transfer and analysis is fully automated in such a way that a naive user can, in theory, start the system with a mouse click and get back the growth analysis results at the end of the lifetime of the plant with no intermediate intervention. As evidence of its functionality, we show and analyze quantitative results of the rhythmic growth patterns of the dicot Arabidopsis thaliana(L.), and the monocot barley (Hordeum vulgare L.) plants under their diurnal light/dark cycles.
Why it matches plant phenotyping methods3Dレーザースキャン、ロボット自動取得、3D再構成により植物の表面積・体積・成長を抽出する自動フェノタイピングシステムが研究の中心である。
titleMachine Vision System for 3D Plant Phenotyping
Thin leaves, fine stems, self-occlusion, non-rigid and slowly changing structures make plants difficult for three-dimensional (3D) scanning and reconstruction -- two critical steps in automated visual phenotyping. Many current solutions such as laser scanning, structured light, and multiview stereo can struggle to acquire usable 3D models because of limitations in scanning resolution and calibration accuracy. In response, we have developed a fast, low-cost, 3D scanning platform to image plants on a rotating stage with two tilting DSLR cameras centred on the plant. This uses new methods of camera calibration and background removal to achieve high-accuracy 3D reconstruction. We assessed the system's accuracy using a 3D visual hull reconstruction algorithm applied on 2 plastic models of dicotyledonous plants, 2 sorghum plants and 2 wheat plants across different sets of tilt angles. Scan times ranged from 3 minutes (to capture 72 images using 2 tilt angles), to 30 minutes (to capture 360 images using 10 tilt angles). The leaf lengths, widths, areas and perimeters of the plastic models were measured manually and compared to measurements from the scanning system: results were within 3-4% of each other. The 3D reconstructions obtained with the scanning system show excellent geometric agreement with all six plant specimens, even plants with thin leaves and fine stems.
Why it matches plant phenotyping methods植物の3D形状・葉形質を取得するスキャン基盤を開発し、実植物および模型で精度検証しており、フェノタイピング手法が研究の中心です。
abstractwe have developed a fast, low-cost, 3D scanning platform to image plants on a rotating stage with two tilting DSLR cameras centred on the plant.
Airborne LiDAR point cloud representing a forest contains 3D data, from which vertical stand structure even of understory layers can be derived. This paper presents a tree segmentation approach for multi-story stands that stratifies the point cloud to canopy layers and segments individual tree crowns within each layer using a digital surface model based tree segmentation method. The novelty of the approach is the stratification procedure that separates the point cloud to an overstory and multiple understory tree canopy layers by analyzing vertical distributions of LiDAR points within overlapping locales. The procedure does not make a priori assumptions about the shape and size of the tree crowns and can, independent of the tree segmentation method, be utilized to vertically stratify tree crowns of forest canopies. We applied the proposed approach to the University of Kentucky Robinson Forest - a natural deciduous forest with complex and highly variable terrain and vegetation structure. The segmentation results showed that using the stratification procedure strongly improved detecting understory trees (from 46% to 68%) at the cost of introducing a fair number of over-segmented understory trees (increased from 1% to 16%), while barely affecting the overall segmentation quality of overstory trees. Results of vertical stratification of the canopy showed that the point density of understory canopy layers were suboptimal for performing a reasonable tree segmentation, suggesting that acquiring denser LiDAR point clouds would allow more improvements in segmenting understory trees. As shown by inspecting correlations of the results with forest structure, the segmentation approach is applicable to a variety of forest types.
Why it matches plant phenotyping methods航空LiDAR点群から森林の個体樹冠と林冠層を抽出・分割する手法を開発し、検出率や過分割率で技術評価している。樹木の構造・配置という植物状態の取得が中心であり、単なる森林測定ではない。
abstractThis paper presents a tree segmentation approach for multi-story stands that stratifies the point cloud to canopy layers and segments individual tree crowns within each layer using a digital surface model based tree segmentation method.
Autonomous crop monitoring at high spatial and temporal resolution is a critical problem in precision agriculture. While Structure from Motion and Multi-View Stereo algorithms can finely reconstruct the 3D structure of a field with low-cost image sensors, these algorithms fail to capture the dynamic nature of continuously growing crops. In this paper we propose a 4D reconstruction approach to crop monitoring, which employs a spatio-temporal model of dynamic scenes that is useful for precision agriculture applications. Additionally, we provide a robust data association algorithm to address the problem of large appearance changes due to scenes being viewed from different angles at different points in time, which is critical to achieving 4D reconstruction. Finally, we collected a high quality dataset with ground truth statistics to evaluate the performance of our method. We demonstrate that our 4D reconstruction approach provides models that are qualitatively correct with respect to visual appearance and quantitatively accurate when measured against the ground truth geometric properties of the monitored crops.
Why it matches plant phenotyping methods作物の動的3D構造を時空間的に再構成する手法を開発し、作物の幾何学的特性を用いて定量評価している。高品質データセットと正解統計による検証も含み、表現型取得が中心である。
abstractIn this paper we propose a 4D reconstruction approach to crop monitoring, which employs a spatio-temporal model of dynamic scenes
We propose a robust method for estimating dynamic 3D curvilinear branching structure from monocular images. While 3D reconstruction from images has been widely studied, estimating thin structure has received less attention. This problem becomes more challenging in the presence of camera error, scene motion, and a constraint that curves are attached in a branching structure. We propose a new general-purpose prior, a branching Gaussian processes (BGP), that models spatial smoothness and temporal dynamics of curves while enforcing attachment between them. We apply this prior to fit 3D trees directly to image data, using an efficient scheme for approximate inference based on expectation propagation. The BGP prior's Gaussian form allows us to approximately marginalize over 3D trees with a given model structure, enabling principled comparison between tree models with varying complexity. We test our approach on a novel multi-view dataset depicting plants with known 3D structures and topologies undergoing small nonrigid motion. Our method outperforms a state-of-the-art 3D reconstruction method designed for non-moving thin structure. We evaluate under several common measures, and we propose a new measure for reconstructions of branching multi-part 3D scenes under motion.
Why it matches plant phenotyping methods植物の3D分枝構造を画像から推定する手法を開発し、植物データセットで既存手法との比較評価と新たな再構成指標の提案を行っており、フェノタイピング手法が中心である。
abstractWe propose a robust method for estimating dynamic 3D curvilinear branching structure from monocular images.