RiceMultimodalX-ray / CTRootMorphology / geometry measurementPhysiological trait estimationGrowth / time-series analysisGrowth / development / phenologyRoot system architecture
Rhizosphere oxidation is a key adaptive mechanism in reductive soil environments, in which oxygen released from roots alters rhizosphere redox conditions and regulates biogeochemical processes. Rice plants possess an internal oxygen transport system, and radial oxygen loss (ROL) from roots is closely associated with root development. However, the spatial patterns of ROL in soil and their relationships with root traits remain poorly characterized. In this study, we developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice. Daily time-course tracking of individual crown roots revealed dynamic changes in the spatial distribution and magnitude of rhizosphere oxygen in relation to root elongation and aging. Root thickness was positively correlated with dissolved oxygen levels near root tips. Genotypic comparisons further identified a cultivar with reduced rhizosphere oxidation despite possessing thicker roots among the tested genotypes, thereby indicating the involvement of additional physiological processes. Overall, these findings demonstrate that rhizosphere oxidation is regulated by root growth stage and thickness and dynamically modulated during root development.
Why it matches plant phenotyping methods平面酸素オプトードとX線CTを統合したマルチモーダル画像システムを開発し、根の発達と根圏酸化を時系列・個体別に定量化しており、表現型取得手法が研究の中心である。
abstractwe developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice.
Reproduction assets foundThe paper's Data availability statement explicitly deposits the authors' RG2DO-Root analysis program together with sample optode and CT images (the paper's phenotyping inputs) in a public GitHub repository, matching the allowed URL.Code · publicing
8
This work was supported by project JPNP18016, commissioned by the New Energy and
9
Industrial Technology Development Organization (NEDO), JST CREST (JPMJCR17O1),
10
and JST ALCA-Next (JPMJAN23D3).
11
12
Data availability
13
The source code and sample data (optode and CT images) are available from the
14
GitHub repository (https://github.com/tsubasa-kawai28/RG2DO-Root).15
16
References
17
Aguilar EA et al. 2003. Oxygen distribution and movement, respiration and nutrient
18
loading in banana roots (Musa spp. L.) subjected to aerated and oxygen-depleted
19
environments. Plant Soil. 253:91–102. https://doi.org/10.1023/A:1024598319404.20
Armstrong W, Wright EJ. 1975. Radial oxygen loss fromOpen asset ↗https://github.com/tsubasa-kawai28/RG2DO-Root · RG2DO-Rootpdf-raw-page:19 lines:1-82Code / dataset availability confirmedEurope PMC · Crossref · checked 15 Sept 2026
Published3 Sept 2026Methods in ecology and evolution
Growth chamberMultimodalStereoRootStem / branchTrackingGrowth / development / phenology
Understanding plant behaviour requires the integration of multiple phenotypic and physiological signals measured over time under controlled conditions. However, different plant signals are typically studied using separate experimental setups, limiting temporal alignment and integrative analyses.We present Mind(the)Plant, a modular experimental facility designed for the synchronized, long-term acquisition of multimodal plant data, including three-dimensional shoot kinematics, above- and below-ground volatile organic compounds (VOCs) and root imaging. Its modular architecture is designed to accommodate additional acquisition modules, such as electrophysiological signalling, as future extensions. The platform integrates a controlled growth environment with stereovision imaging, high-resolution time-of-flight mass spectrometry and custom rhizocameras. These components are connected through a unified network infrastructure that ensures synchronized acquisition and centralized data handling.We validate the performance of each acquisition module through multi-week recordings, demonstrating high-temporal stability, reliable stereovision synchronization, effective isolation of VOCs signals and robust operation of below-ground imaging. We further illustrate the analytical potential of the platform using a one-day continuous multimodal acquisition combining shoot kinematics, above-ground VOC emissions, rhizocameras observations and environmental data.Mind(the)Plant provides a novel methodological framework for studying plant behaviour, signalling and phenotypic plasticity in ecological and evolutionary research. By enabling coordinated measurements of multiple plant response modalities, the platform supports investigations of dynamic plant-environment and plant-plant interactions from a behavioural perspective.
Why it matches plant phenotyping methods植物の複数の表現型・生理シグナルを同期取得する施設を開発し、各取得モジュールの性能を検証しているため、表現型計測プラットフォームが研究の中心です。
abstractWe present Mind(the)Plant, a modular experimental facility designed for the synchronized, long-term acquisition of multimodal plant data, including three-dimensional shoot kinematics, above- and below-ground volatile organic compounds (VOCs) and root imaging.
Reproduction assets foundThe paper's data availability statement explicitly deposits data, code and processing pipelines (supporting the multimodal plant phenotyping measurements and analysis) in a public Zenodo archive with an authors' URL matching an allowed URL.Code · publicf Interest Statement
The authors have no conflicts of interest to declare.
Peer Review
The peer review history for this article is available at https://www.webofscience.com/api/gateway/wos/peer-review/10.1111/2041-210x.70411 .
Data availability Statement
Data, code and processing pipelines supporting this study are available at https://doi.org/10.5281/zenodo.22095454 ( Simonetti & Castiello, 2026 ).
References
Avesani S, Bonato B, Simonetti V, Guerra S, Ravazzolo L, Gjinaj G, Dadda M, Castiello U. Comparing proton transfer reaction (PTR) and adduct ionization mechanism (AIM) for the study of volatile organic compounds. Molecules. 2026;31(3):402. doi: 10.3390/molecules31030402.
Baluška F, LeOpen asset ↗zenodo · 10.5281/zenodo.22095454lines:482-508Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
Published1 Sept 2026International Journal of Applied Earth Observation and Geoinformation
Accurate plant disease detection remains challenging when using single-modality data, which fails to capture comprehensive disease-related features. However, many existing studies rely on pixel-level classification or prior plant segmentation and lack explicit modeling of cross-modal interactions, limiting their ability to distinguish between healthy and diseased plants. This study presents a spatial–spectral fusion deep learning framework (S2-PDD) that fuses uncrewed aerial vehicles (UAV)-based RGB imagery and hyperspectral-derived feature representations for plant-level detection of potato diseases (blackleg and Potato virus Y (PVY)), employing early fusion (modalities combined at the input stage) and middle fusion (features integrated at intermediate stages within the model backbone) strategies. The multimodal fusion models were compared against single-modal models and an existing S2ADet model. Model performance, assessed through five-fold cross-validation, demonstrated that multimodal models integrating RGB and vegetation index features achieved the highest mAPs of 86.65 ± 1.71 (%; E-RV model) and 85.74 ± 1.96 (M-RV), respectively. These mAPs were higher than those of all single-modal models, including the RGB-only (83.21 ± 1.46) and hyperspectral-only models (PCA features: 79.71 ± 1.45; vegetation index features: 85.31 ± 2.36). They also exceeded mAPs of multimodal models combining RGB with PCA features (early fusion: 83.00 ± 2.81; middle fusion: 83.11 ± 2.46; S2ADet: 84.04 ± 2.73), regardless of the fusion strategy. The superior performance highlights that vegetation index features provide strong class separability compared to other hyperspectral representations. The proposed models achieved strong plant-level detection performance, with AP of 78.65 ± 4.14 (E-RV) and 77.40 ± 3.46 (M-RV) for blackleg disease, as well as 84.82 ± 3.59 (E-RV) and 83.28 ± 3.86 (M-RV) for PVY. These results demonstrate the potential of UAV-based multimodal sensing for disease monitoring in cropping systems. A potato plant disease detection dataset was constructed and made publicly available, containing paired RGB and hyperspectral image tiles with bounding box annotations. The code is available at https://github.com/Tim-Agro/S2-PDD.
Why it matches plant phenotyping methodsUAV RGB・ハイパースペクトル画像からジャガイモ個体の病害状態を推定する融合モデルを開発・比較検証し、公開データセットも構築しており、病害表現型の取得・抽出手法が中心である。
abstractThis study presents a spatial–spectral fusion deep learning framework (S2-PDD) that fuses uncrewed aerial vehicles (UAV)-based RGB imagery and hyperspectral-derived feature representations for plant-level detection of potato diseases (blackleg and Potato virus Y (PVY))
Plant phenotyping relevance match · UnverifiedOpenAlex · Crossref · Europe PMC · checked 14 Sept 2026
Crop phenotyping is crucial for advancing plant breeding, yet remains a significant challenge. Manual approaches are labor-intensive and do not scale to the analysis of large datasets, while computational methods like Vision-Language Models (VLMs) lack the adaptability for fine-scale spatial reasoning and diverse phenotyping scenarios. To bridge the gaps, we present iPheno, a fully open-source, domain-specialized multimodal VLM for fine-scale crop phenotyping. A key innovation of iPheno is its dual-aware architecture. First, a spatial-aware feature extractor samples mask regions into local K-Nearest Neighbors (KNN) graphs and enables fine-scale analysis of arbitrary-shaped regions; second, a task-aware Mixture-of-Experts (MoE) routing mechanism activates specialized modules for each phenotyping task. To train iPheno and achieve rigorous benchmarking, we constructed iPheno-120K, a large-scale high-precision dataset designed for multiple phenotyping tasks. Evaluations on iPheno-120k test set and other publicly available datasets showed that iPheno outperformed all fine-tuned baselines, by improving F1-score by 17.1% (LLaVA-1.6-13B) to 28.9% (MiniCPM-o-9B), while achieving the highest inference speed and memory efficiency. A web server (https://ipheno.ai4bread.com), a mobile application (www.ipheno.cn), and a stand-alone PC client (https://github.com/2997029323/iPheno-PC-Client) are available for iPheno.
Why it matches plant phenotyping methods植物表現型を対象とするVLMの開発、マルチタスク評価、専用データセット構築が研究の中心であり、明確な方法論的貢献がある。
abstractwe present iPheno, a fully open-source, domain-specialized multimodal VLM for fine-scale crop phenotyping.
Background Understanding the structure of plant seeds cultivated for human consumption and food manufacturing is vital to provide sustainable products as well as to investigate early growth stages. This includes structural variation between different plant species, varieties and cultivars depending on genetic setup, as well as structural modifications upon germination, aging and storing or seed treatment during processing. For plant seeds as multi-component biological materials, structural characterization must extend across multiple length scales, from molecular organization to cellular architecture. Results We apply scanning Small- and Wide-Angle X-ray Scattering (SWAXS) and X-ray Fluorescence (XRF) on yellow pea seeds to combine local structural information on the molecular scale with imaging of cellular structures on the micrometer scale, enabling a comprehensive analysis of hierarchical organization. To identify and characterize heterogeneous regions within the pea seeds, we implement a fitting-free, data-driven segmentation and analysis workflow based on machine learning tools. This approach allows for classification of structurally distinct domains and enables quantitative comparison across samples without relying on predefined models. Furthermore, we incorporate multi-modal analysis by combining structural imaging with complementary elemental information obtained from XRF. The integration of compositional and structural data provides deeper insight into structure-composition relationships. Conclusions This multi-scale, multi-modal approach opens new possibilities for investigating hierarchical structures and their development under diverse conditions and enables systematic comparison between different species or seeds at different developmental stages or exposed to different processing steps. The approach is broadly applicable to various kinds of samples and other hierarchically organized biological materials, which makes it a valuable technique for plant science as well as plant-based food science.
Why it matches plant phenotyping methods種子の構造・細胞領域をX線散乱/蛍光イメージングと機械学習ベースのセグメンテーションで定量解析する手法が研究の中心であり、植物器官の構造形質を抽出するため採用。
abstractWe apply scanning Small- and Wide-Angle X-ray Scattering (SWAXS) and X-ray Fluorescence (XRF) on yellow pea seeds to combine local structural information on the molecular scale with imaging of cellular structures on the micrometer scale
Plant phenotyping relevance match · UnverifiedOpenAlex · Europe PMC · Crossref · bioRxiv · checked 5 Sept 2026
Abstract Non-invasive, high-throughput phenotyping tools are needed that can identify environmental effects on plant structure and function to diagnose factors responsible for reduced growth in commercial and non-commercial settings. In this study, we explored whether the integration of 3D-multispectral (3D) and 2D-hyperspectral imaging (HSI), aided by machine learning (ML), could be used to identify environmental stress treatments imposed during plant growth. Controlled environment-grown Nicotiana Benthamiana plants were subjected to a range of abiotic treatments – including different growth irradiances, heat treatment and drought stress – with the treatments resulting in differences in shoot height, biomass, leaf area and spectral reflectance. ML models were trained to identify these treatments using morphological and spectral traits measured at 27, 29, 31, and 34 days after sowing (DAS). A 3D-multispectral scanner was used to obtain information on plant height, biomass, and leaf area. A visible and near-infrared (VNIR) HSI camera provided detailed spectral information for deriving spectral indices including the Normalised Difference Vegetation Index (NDVI), Photochemical Reflectance Index (PRI) and Normalized Difference Red Edge (NDRE). Manual measurements provided baseline comparative data. The 3D-multispectral scanner reliably estimated above-ground traits, with high correlations between manual and scanner-derived measurements. The ML models accurately differentiated among environmental stress treatments, with the fused 3D+HSI model achieving the best overall predictive performance across all evaluated metrics compared with models based on either imaging modality alone. Results demonstrated the effectiveness of combining 3D-multispectral and 2D-HSI data with ML analyses for non-destructive, high-throughput phenotyping. The integration of these techniques enabled non-destructive, high-throughput identification of environmental stress treatments imposed during plant growth.
Why it matches plant phenotyping methods3Dマルチスペクトル画像・ハイパースペクトル画像と機械学習を統合し、植物形態・スペクトル形質を非破壊かつ高スループットに取得・検証する方法が研究の中心である。
abstractNon-invasive, high-throughput phenotyping tools are needed that can identify environmental effects on plant structure and function
Accurate quantification of rice aboveground biomass (AGB) is critical for crop monitoring but remains challenging due to the complex nonlinearity arising from the coupling of plant density, spatial structure, and internal dry matter distribution. To address the limitations of single-source remote sensing, this study proposed a Three-Dimensional Dry Matter Distribution Integration (3D-DMI) model, which establishes a physically interpretable framework decomposing AGB into dry matter density ( ρ ), horizontal projection distribution ( S ), and vertical cumulative distribution ( h d ) components. Guided by this framework, a core subset of six features (Red_650, MTCI, G_correlation, R_correlation, LPI, and HPA0_99) was extracted from UAV-based multispectral, RGB, and LiDAR data using a dual-step feature selection approach combining Maximum Information Coefficient (MIC) and Distance Correlation (dCor). A Random Forest (RF) regression model was then developed to estimate AGB across the entire growth season. The results demonstrated that the 3D-DMI model achieved excellent performance with an R 2 of 0.920, an RMSE of 0.184 kg/m², and an RPD of 3.544, significantly outperforming any single-sensor approach. Single-feature analysis revealed that while LiDAR-derived structural features provided the fundamental basis for biomass estimation, they encountered inherent saturation bottlenecks during late growth stages. Feature contribution analysis based on SHAP further quantified that LiDAR-derived features dominated the estimation process (68.5% contribution), providing the volumetric basis, whereas RGB textures (18.3%) and multispectral features (13.3%) provided indispensable supplements. Ultimately, this study established a robust, physically grounded computational paradigm for high-precision UAV-based rice biomass monitoring across the entire growth cycle.
Why it matches plant phenotyping methodsUAVマルチスペクトル・RGB・LiDARデータからイネの地上部バイオマスを推定する3D-DMI計算手法を開発・評価しており、植物形質の取得・抽出が研究の中心である。
abstractthis study proposed a Three-Dimensional Dry Matter Distribution Integration (3D-DMI) model
Herbarium specimens are physical, verifiable records that form the basis of taxonomic knowledge and biodiversity research. Their large-scale digitization has produced extensive collections of high-resolution images and associated specimen metadata, creating conditions in which artificial intelligence (AI) can play an important role in plant taxonomy, collection management, and ecological research. Early AI applications have primarily focused on automated species identification based on individual specimen images. Although increasingly accurate, such approaches remain limited by their emphasis on single-specimen label prediction and by treating identification outputs as final analytical decisions. Recent methodological advances-including segmentation-based preprocessing, automated trait extraction, structured extraction of label data, detection of potentially misidentified specimens, and multimodal integration of visual, textual, and genetic information-extend AI applications beyond species identification toward broader analytical frameworks, encompassing taxonomic interpretation as well as ecological and biodiversity research. In these approaches, specimens are placed within a shared analytical space, and identification results are used to support comparisons across multiple specimens rather than being treated as final decisions for single individuals. This multi-specimen perspective enables quantitative examination of species boundaries, morphological variation, data inconsistencies, and taxonomic stability within curated collections. In this context, AI serves not as an ultimate decision-maker but as a decision-support tool embedded in expert-guided workflows and biodiversity knowledge infrastructures. These developments can be summarized as Integrative Taxonomic AI, an approach that employs learned morphospaces to interpret and refine taxonomic categories by integrating multimodal evidence and curated specimen data under expert guidance.
Why it matches plant phenotyping methods植物標本画像から形態形質を抽出するAI手法と統合的解析枠組みを中心に扱うレビューであり、植物表現型取得・抽出法との関連が明確。
abstractRecent methodological advances-including segmentation-based preprocessing, automated trait extraction, structured extraction of label data, detection of potentially misidentified specimens, and multimodal integration of visual, textual, and genetic information-extend AI applications beyond species identification
Precision agriculture increasingly requires intelligent systems capable of integrating multimodal sensing with transparent decision support to enable timely and reliable crop management. This study proposes a hybrid intelligent IoT framework integrating environmental monitoring, wearable plant physiological sensing, AI-based pest-monitoring, machine-learning-based prediction of crop physiological stress, and explainable fuzzy rule-based decision support into a unified architecture for crop stress assessment. A novel Physiological Stress Index (PSI) was developed by combining vapor pressure deficit, relative humidity, Delta-T, leaf capacitance, and relative irradiance to provide an interpretable indicator of crop physiological stress. The proposed framework was experimentally validated under real field conditions using a commercial environmental monitoring station, wearable leaf sensors, AI-enabled pest-monitoring devices, and cloud-based analytics. Correlation analysis confirmed strong relationships between PSI and the principal environmental variables (VPD: r = 0.980, Delta-T: r = 0.990, RH: r = −0.961), demonstrating the internal consistency and sensitivity of the proposed index. At the 15 min forecasting horizon, Linear Regression and Gradient Boosting demonstrated virtually identical performance: Gradient Boosting achieved a marginally lower RMSE and higher R2 (RMSE = 0.0273; R2 = 0.9810), whereas Linear Regression achieved a slightly lower MAE (MAE = 0.0186). At the 1 h forecasting horizon, Gradient Boosting achieved the strongest performance (R2 = 0.9034), indicating increasing relevance of nonlinear modelling at longer prediction horizons. The proposed framework demonstrates the feasibility of combining multimodal sensing, machine learning, explainable artificial intelligence, and edge-enabled IoT technologies to support proactive, transparent, and intelligent precision agriculture.
Why it matches plant phenotyping methods植物の生理的ストレス状態を多モーダルセンサーと機械学習で推定する方法を開発し、圃場で検証しており、フェノタイピング手法が中心である。
Plant diseases destroy 20–40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.
Why it matches plant phenotyping methods植物病害の画像ベース検出・予測手法を対象とする系統的レビューであり、植物の病徴・病害状態を観測から推定するフェノタイピング手法のレビューとして中心的です。
titleMultimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases
Abstract In real field scenarios in agriculture, automatic segmentation of plant diseases is an important technique for precision farming. However, it remains exceptionally challenging due to blurred lesions, complex morphological structures, irregular backgrounds, and severe class imbalance. While traditional convo lutional networks struggle to capture long-range semantic context and standard vision transformers fail to preserve sharp localized boundaries, this paper proposes an efficient, attention-gated hybrid framework optimized for field deployment. Our architecture leverages a hierarchical Mix Transformer (MiT-B2) encoder stream integrated with an Atrous Spatial Pyramid Pooling (ASPP) scale-space context bridge and a custom Cross-Scale Multimodal Attention Gate (CMAG) to isolate discriminative disease markers selectively. Evaluated on the highly challenging and unbalanced PlantSeg dataset, our framework achieves competitive mean Intersection over Union (mIoU) of 66.57% and an F1-score of 79.93%, while maintaining a highly compact parameter footprint of only 30.37 M. Experimental evaluations demonstrate that the proposed system establishes a new performance milestone, outperforming current competitive architectures and proving highly viable for resource-constrained edge devices. To further enhance out-of-distribution stability, we outline future directions to extend our top-performing candidate variants into a Level 1 meta-stacking ensemble optimized via few-shot learning and partial backbone fine-tuning.
Why it matches plant phenotyping methods植物病害領域の画像セグメンテーション手法を開発・評価し、病変の分割性能を定量検証しているため、植物の病害状態を推定するフェノタイピング手法が中心です。
abstractautomatic segmentation of plant diseases is an important technique for precision farming.
Plant phenotyping is essential for modern crop breeding, yet traditional static image analysis fails to capture the nonlinear dynamics of plant growth. Existing time-series forecasting models exhibit notable limitations when processing multimodal data: global pooling operations may compress local 2D spatial topology of plants, and shallow feature concatenation may be insufficient for effective cross-modal semantic alignment. Moreover, current methods typically regress absolute morphological states, which may contribute to temporal lag during nonlinear growth spurts. In this paper, we propose ST-CrossGro-Former, a cross-modal residual forecasting framework for plant dynamic growth. The network removes the final global pooling and classification layers to preserve spatial topology and incorporates a scalar-guided cross-modal attention module based on the standard query-key-value formulation. This module utilizes 1D morphological features as queries to dynamically weight local visual regions, promoting multimodal feature alignment. Concurrently, a residual incremental forecasting strategy is introduced to predict short-term growth increments rather than absolute states, aiming to improve tracking sensitivity to sudden growth events. Evaluations on the UNL-CPPD maize dataset and supplementary validation on the FIP1 wheat field dataset show that the proposed model achieves competitive single-step forecasting accuracy and favorable temporal trajectory alignment compared with adapted spatiotemporal attention, graph-based, and physics-informed baselines under the evaluated settings. In particular, the FIP1 results suggest that ST-CrossGro-Former can maintain favorable height trajectory alignment under a field-acquired wheat setting, indicating its potential for helping mitigate temporal misalignment in dynamic growth forecasting.
Why it matches plant phenotyping methods植物の動的形態成長を予測する新規クロスモーダル手法を開発し、トウモロコシ・コムギデータセットで評価しており、表現型の抽出・予測手法が中心である。
abstractwe propose ST-CrossGro-Former, a cross-modal residual forecasting framework for plant dynamic growth.
Reproduction assets foundThe paper evaluates its ST-CrossGro-Former model on two public plant phenotyping datasets: the UNL-CPPD maize dataset (explicitly stated as publicly available with a repository URL) and the FIP1 wheat field dataset (public dataset from ETH Zürich, with its GigaScience dataset publication DOI). No author analysis code, Dataset · publicThe UNL-CPPD dataset used in this research was acquired from the UNL Plant Phenotyping Datasets repository, accessible at https://plantvision.unl.edu/datasets.Open asset ↗UNL Plant Phenotyping Datasets · UNL-CPPDlines:266-273Plant phenotyping relevance match · UnverifiedOpenAlex · checked 5 Sept 2026
Drought stress severely limits foxtail millet yield and quality, yet current drought-resistance indices are exclusively yield-oriented and ignore grain-filling quality. Our two-year (2024–2025) experiments with 24–48 varieties revealed that yield and blighted grain rate (BGR) are partially decoupled (e.g., Zhangzagu 18: yield 2307 kg/ha, BGR 0.444; Zhonggu 19: yield 1622 kg/ha, BGR 0.280). We therefore constructed the Yield–Quality Synergy Index (YQSI = DYI − BGR), which penalizes varieties with poor grain filling. The YQSI tied for first place with DYI in comprehensive screening performance and achieved the highest inter-annual stability (Spearman ρ = 0.823, Jaccard = 0.438, composite score = 1.261). Sensitivity analysis confirmed robustness of the equal-weight formula across a 4-fold range of quality-penalty weights. Six strongly drought-resistant germplasms with balanced yield and quality were identified. Using UAV multimodal data (RGB, multispectral, and thermal infrared) acquired during grain filling, a Random Forest model predicted a YQSI with overall R2 = 0.819 and an F1 score of 0.933 for variety screening. Feature-importance analysis highlighted NDVI, WDRVI, and red-edge texture as key predictors. This study provides a quality-constrained drought-resistance evaluation framework and demonstrates the potential of UAV-based high-throughput phenotyping for foxtail millet breeding.
Why it matches plant phenotyping methodsUAVのRGB・マルチスペクトル・熱赤外データから干ばつ耐性指標を予測する高スループット表現型解析手法が研究の中心であり、モデル性能も評価している。
abstractUsing UAV multimodal data (RGB, multispectral, and thermal infrared) acquired during grain filling, a Random Forest model predicted a YQSI with overall R2 = 0.819 and an F1 score of 0.933 for variety screening.
Sustainable wheat farming is challenging. Real-time information on crop health, disease transmission, and anticipated yields is essential for farmers. However, they frequently use slow, expensive, or non-communicative tools. This project develops a workable solution. There is no need for massive server farms because the entire system operates on a single graphics card. It incorporates images of wheat fields, Indian farming notes, greenhouse records, harvest statistics, and NASA meteorological data. Consider them as various “eyes” for crop photo analysis, and we tried several lightweight computer vision models. ConvNeXt-Tiny was slower but could operate on older equipment with 75% accuracy; EfficientNetB0 recognised wheat heads with 92% accuracy; and AgroMark, a hybrid solution that merged photo analysis with agricultural metadata (soil type, rainfall, increased to 87%, etc. Combining picture analysis with attention mechanisms (CBAM) allowed us to anticipate the amount of wheat that a field will yield based on these photo insights, and the results showed that our predictions were accurate, with an R 2 score of 0.97. Additionally, we developed a versatile detector that simultaneously detects disease, stress, head count, and pests. It is adjusted to deal with training data that is unbalanced (some diseases are common, while others are rare). As we packed everything into a 16-GB graphics card, we spent real time determining which strategies smaller training sets, removing weak features, and adjusting loss functions, work. We encounter real-world obstacles along the road, such as photographs from different locations not always match, mislabeled photographs from different locations not always match, mislabeled diseases, and neglected rare pests. Our step-by-step instructions, charts, and code are available.
Why it matches plant phenotyping methods小麦画像から病害・ストレス・穂数・収量などの植物形質・状態を推定するマルチモーダル手法を開発し、複数モデルの精度比較と実装上の検証を行っており、表現型取得・推定が研究の中心である。
abstractThis project develops a workable solution.
Reproduction assets foundThe paper builds its multimodal wheat phenotyping analysis on several explicitly cited public data assets: the Kaggle Wheat Plant Diseases image dataset (used for disease classification, Tables 2 and 9), the Global Wheat Head Detection dataset (used for head detection, Tables 1 and 6), FAOSTAT and India Open GovernmentDataset · publicAvailable online at: https://www.fao.org/faostat/ . FAOSTAT statistical database.Open asset ↗lines:1110-1162Plant phenotyping relevance match · UnverifiedOpenAlex · Crossref · checked 15 Sept 2026
Field / plotMultimodalNeRF / 3D Gaussian SplattingLiDAR / point cloudRGB / grayscaleStereoFruitWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detection
Binocular stereo vision is a low-cost and scalable 3D perception technology that shows strong potential in agricultural phenotyping and smart agriculture. By estimating depth from multi-view RGB images, it enables non-contact, high-precision sensing of crop structure, canopy morphology, growth dynamics, and livestock traits, providing essential support for digital and intelligent agricultural production. With recent advances in deep learning-based stereo matching, multimodal sensor fusion, and 3D reconstruction, its robustness and accuracy in complex field environments have been significantly improved. This paper systematically reviews recent progress in agricultural applications of binocular stereo vision, covering system architectures, traditional and deep learning-based stereo matching methods, point cloud reconstruction techniques, and emerging supervision strategies such as 3D Gaussian splatting. It further summarizes key applications, including high-throughput phenotyping, fruit localization and robotic harvesting, weed detection and precision spraying, autonomous navigation, and livestock body condition assessment, highlighting its role in multi-task agricultural perception systems. Finally, the paper discusses major challenges, including low-texture matching difficulty, occlusions in complex environments, cross-domain generalization, real-time lightweight deployment, and limited dataset availability. Future directions are outlined in foundation model-based visual perception, self- and weakly supervised learning, multimodal fusion, and edge-efficient model design, aiming to support large-scale deployment in smart agriculture.
Why it matches plant phenotyping methods農業における双眼ステレオビジョンのシステム、ステレオマッチング、3D再構成を体系的にレビューし、作物構造・群落形態・生育動態の非接触計測とハイスループット表現型解析を主要対象としているため。
abstractThis paper systematically reviews recent progress in agricultural applications of binocular stereo vision, covering system architectures, traditional and deep learning-based stereo matching methods, point cloud reconstruction techniques
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
The use of glufosinate-resistant GM soybean has expanded, raising concerns about resistant weed development and unintended transgene flow. To support monitoring for timely management, we propose an early, non-destructive identification method using spectral images acquired from whole soybean plants after glufosinate treatment. We evaluated the potential of spectral imaging, using RGB, infrared (IR) thermal, and chlorophyll fluorescence (CF) sensors, for early detection of glufosinate resistance in soybean. In the dose-response test, the key spectral indices including NDI, temperature difference, F v /F m , and NPQ distinguished between resistant and susceptible soybeans within 4 to 24 hours after treatment (HAT). IR thermal and CF imaging showed higher sensitivity in identifying resistance than RGB imaging by detecting spectral responses associated with physiological changes before visual symptoms appeared. Validation test with a single dose treatment of glufosinate reconfirmed that image analysis by both the naked eye and machine learning (ML) can discriminate between resistant and susceptible soybeans in a single day after glufosinate treatment. ML-based classification using IR thermal index achieved 100% accuracy as early as 6 HAT and the classification by the naked eye using IR thermal images showed 96.6% accuracy at 24 HAT. These results suggest that plant imaging enables early and non-destructive identification of herbicide-resistant individuals by detecting early spectral changes to herbicide treatment. These findings support its use as a potential alternative to conventional diagnostic methods for detecting individuals containing transgenes in herbicide-resistant GM soybean cultivation for future applications in herbicide-resistant weed monitoring.
Why it matches plant phenotyping methodsスペクトル画像(RGB、熱赤外、クロロフィル蛍光)と機械学習を用いて、薬剤処理後の植物の生理応答から耐性を早期識別する方法を開発・検証しており、植物表現型の取得が中心である。
abstractwe propose an early, non-destructive identification method using spectral images acquired from whole soybean plants after glufosinate treatment.
Traditional breeding is approaching its efficiency limit in addressing global food security challenges, climate change, and multi-dimensional information integration. Artificial intelligence (AI) and digital phenotyping have become core driving forces in data-driven breeding. This review systematically elaborates the evolutionary roadmap from AI breeding 1.0 , which relies on traditional phenotypic data, to the emerging paradigm of AI breeding 2.0 . We clarify the core differences between these two paradigms, dissect the “data divide” in the transition process, and summarize the integrated technical framework underlying AI breeding 2.0 . The central argument is that the core limitation of AI breeding 1.0 lies in data inadequacy rather than constraints on algorithmic capacity. The transition to AI breeding 2.0 relies on standardized high-throughput digital phenotyping, multi-modal data integration, and closed-loop data-centric breeding systems. Finally, we propose four priority actions to promote the widespread adoption of data-driven breeding and provide an actionable roadmap for enhancing global crop improvement efforts.
Why it matches plant phenotyping methods植物デジタルフェノタイピングをAI育種の中核技術として体系的に論じるレビューであり、フェノタイピング手法・データ統合・技術枠組みが中心です。
abstractArtificial intelligence (AI) and digital phenotyping have become core driving forces in data-driven breeding.
Growth chamberMultimodalVisualization / data management
Indoor high-throughput plant phenotyping (HTPP) platforms require flexible, interoperable architectures to support reproducible trait acquisition across changing controlled-environment experiments. This paper presents GREENTRIBE, a modular multi-sensor indoor HTPP framework that integrates sensing, robotic coordination, data communication, semantic data management, and crop modelling within a unified phenotyping pipeline. Rather than treating these elements as independent modules, GREENTRIBE connects distributed sensor acquisition, robot-assisted operation, lightweight message exchange, ontology-based data organization, and process-based model assimilation through a layered architecture implemented with CAN, ROS 2, MQTT, OpenSILEX, and STICS. The platform combines a multiscale sensing network with a sensor-independent communication protocol, enabling heterogeneous imaging, environmental, and plant-monitoring devices to be configured within indoor experimental designs. To support traceable and reusable workflows, GREENTRIBE implements an ontology-driven data management layer aligned with FAIR (Findable, Accessible, Interoperable, and Reusable) metadata standards. The architecture further links computer vision and artificial intelligence pipelines with the STICS crop model, allowing multimodal observations to be transformed into biologically interpretable phenotypic information under explicit genotype, environment, and management contexts. Platform validation demonstrated reliable communication, efficient multimodal data handling, and standardized metadata management across the sensing-to-information pipeline. Under the configured acquisition schedule, GREENTRIBE achieved a maximum full data cycle of approximately 200 ms, no measurable losses up to the CAN master, high metadata completeness, and support for 1704 scheduled daily acquisition events from seven devices. Overall, GREENTRIBE provides a modular and interoperable indoor phenotyping framework for reproducible experiments and standardized multimodal phenotypic data generation.
Why it matches plant phenotyping methods植物フェノタイピングのための多センサーHTPP基盤を中心に開発・統合し、通信性能、データ処理、メタデータ管理を検証しているため。
abstractThis paper presents GREENTRIBE, a modular multi-sensor indoor HTPP framework that integrates sensing, robotic coordination, data communication, semantic data management, and crop modelling within a unified phenotyping pipeline.
ABSTRACT Timely and reliable assessment of plant health is essential for resilient and sustainable agriculture, yet diagnostic practice remains divided between accurate but laboratory‐bound assays and emerging field‐deployable technologies. Herein, we adopt a plant‐centric and systems‐level perspective to synthesize advances in smart sensing for plant health monitoring reported between 2000 and 2026. Rather than surveying individual devices in isolation, we organize the literature around how sensing technologies are fabricated, integrated, and translated into actionable agronomic insights. Our synthesis reveals a clear shift toward multimodal, minimally invasive sensing architectures that combine material and fabrication innovations with contact and non‐contact modalities spanning organ, canopy, and landscape scales. We find that the most impactful progress arises not from individual sensors alone, but from integrated pipelines that couple sensing hardware with edge intelligence, cross‐scale data fusion, and explainable analytics. Furthermore, persistent barriers, including calibration transfer, long‐term stability, power autonomy, dataset bias, and cybersecurity, continue to impede widespread adoption. Based on these findings, we outline design principles and research priorities needed to accelerate translation, emphasizing standardized validation against biological benchmarks, energy‐autonomous and environmentally responsible sensor systems, and artificial intelligence (AI) frameworks capable of robust generalization across crops and environments. Looking ahead, we argue that plant health monitoring will increasingly be defined by closed‐loop systems that directly link plant physiological or pathological signals to adaptive management, positioning smart sensing as a cornerstone of data‐driven and climate‐resilient agriculture.
Why it matches plant phenotyping methods植物の健康状態・生理・病理シグナルを対象とするスマートセンシング手法を、統合、検証、校正、データ解析の観点から体系的にレビューしており、植物フェノタイピング手法が中心である。
abstractwe outline design principles and research priorities needed to accelerate translation, emphasizing standardized validation against biological benchmarks
Accurate monitoring of cotton plant moisture content (PMC) is crucial for guiding irrigation practices. To address the limited capacity of single-source remote sensing data to characterize the water status of cotton plants, as well as the lack of quantitative reference values for suitable PMC levels at different growth stages, this study constructed a cotton PMC estimation model based on multimodal UAV remote sensing data. Furthermore, the suitable reference levels of PMC at different growth stages were investigated according to the response relationship between PMC and yield at each growth stage. Five soil moisture gradients were established, and at each growth stage, fresh and dry weights of cotton shoots were measured to calculate the PMC. A UAV platform equipped with multiple sensors was used to collect visible-light (RGB), multispectral (MS), and thermal infrared (TIR) images of the cotton canopy. Three feature selection methods were employed to identify moisture-sensitive parameters: Pearson correlation analysis, principal component analysis (PCA) for dimensionality reduction, and recursive feature elimination (RFE). Using the selected parameters, four machine learning algorithms, AdaBoost, random forest (RF), CatBoost, and k-nearest neighbors (KNN), were applied to construct and validate PMC estimation models. The suitable PMC levels at different growth stages were identified based on the response relationship between measured PMC and yield under different water gradients. The results showed that the RFE feature selection method identified eight water-sensitive parameters, and the CatBoost model integrating multimodal data performed best, with R² and RMSE reaching 0.807 and 0.033%, respectively, on the test set, providing a reliable method for high-resolution spatial mapping of field-scale PMC. On this basis, the response of yield to PMC was analyzed, revealing that when PMC was maintained at 83.8%, 85.9%, 79.3%, 78.0%, and 67.7% at the bud, initial flowering, peak flowering, peak boll-setting, and boll opening stages, respectively, the theoretical maximum yield of 6579–6667 kg/hm² could be achieved. This study realized high-precision remote sensing monitoring of PMC and further explored the appropriate moisture content thresholds for different growth stages, providing a quantitative reference for precision water regulation in cotton fields.
Why it matches plant phenotyping methodsUAVのマルチモーダル画像と機械学習により、綿植物の水分含量を推定・検証する手法が研究の中心であり、植物状態の高解像度マッピングにも応用している。
abstractthis study constructed a cotton PMC estimation model based on multimodal UAV remote sensing data
Accurate estimation of forest aboveground biomass (AGB) is critical for carbon accounting, ecosystem monitoring, and climate change mitigation. While remote sensing offers broad spatial coverage, standard deterministic models often struggle to capture the complex relationship between spectral information and structural forest attributes, and they frequently lack reliable quantification of uncertainty. This study presents a novel probabilistic framework that integrates convolutional neural networks (CNNs) and transformer architectures for AGB estimation using a multi-sensor fusion of Sentinel-1/2 imagery. The framework employs a customized, adaptive version of the Swin Transformer-reconfigured specifically for regression tasks and streamlined to a single-stage architecture-and implements a Gaussian Mixture Modeling (GMM) approach to ensemble model outputs. This framework enables the decomposition of predictive uncertainty into epistemic (model-related) and aleatoric (data-inherent) components. Evaluated across heterogeneous forest regions in Montana and Idaho, the CNN-transformer integration achieved superior performance, with R 2 =0.83, RMSE=19.16 Mg/ha, and MAE=13.37 Mg/ha, outperforming standalone CNNs (R 2 =0.81) and vision transformers (R 2 =0.80). Transfer learning and fine-tuning experiments demonstrated high model robustness, improving prediction accuracy in independent test regions by 39% in RMSE. Spatially explicit analysis revealed that while CNNs provide stable local feature extraction, the Swin Transformer’s self-attention mechanism significantly mitigates errors on topographically complex slopes by leveraging non-local spectral cues to compensate for shadowing and geometric distortions. The results highlight that an ensemble approach leveraging the complementary strengths of CNNs and customized transformers provides the transparency and precision required for operational carbon monitoring and high-resolution, trustworthy biomass mapping.
Why it matches plant phenotyping methods森林の地上部バイオマスという植物群落形質を、マルチセンサー画像とCNN・Transformer・GMMで推定する手法を開発し、比較評価と独立地域での検証を行っており、形質推定法が研究の中心である。
abstractThis study presents a novel probabilistic framework that integrates convolutional neural networks (CNNs) and transformer architectures for AGB estimation using a multi-sensor fusion of Sentinel-1/2 imagery.
Abstract. Urban vegetation is essential for mitigating the Urban Heat Island effect, yet its cooling performance depends on its three-dimensional structure. This study combines high-resolution Unmanned Aerial Vehicle - based LiDAR (Zenmuse L2) and thermal imaging (Zenmuse H20) to analyze vegetation structure and surface temperature across 4 urban parks in San Nicolás de los Garza, Mexico. LiDAR data were processed to generate Digital Terrain Model, Digital Surface Model and Canopy Height Model models, enabling the segmentation of individual trees and extraction of structural metrics such as canopy height, crown area and point density. Thermal orthomosaics were co-registered with LiDAR models to quantify temperature contrasts between vegetated and impervious areas. Results reveal consistent cooling effects in all parks, with vegetated zones showing 8–15 °C lower surface temperatures depending on canopy density and maturity. Larger parks with continuous canopies displayed the strongest thermal regulation. This integrated LiDAR–thermal approach provides a precise and scalable framework for assessing microclimatic benefits of urban vegetation, supporting climate-resilient planning in rapidly urbanizing regions.
Why it matches plant phenotyping methodsUAV LiDAR・熱画像を用いて個体樹木の樹冠高や樹冠面積などの植物構造形質を抽出する手法と統合ワークフローが中心であり、単なる環境測定ではない。
abstractLiDAR data were processed to generate Digital Terrain Model, Digital Surface Model and Canopy Height Model models, enabling the segmentation of individual trees and extraction of structural metrics such as canopy height, crown area and point density.
Long-duration flood inundation can substantially suppress crop growth and cause yield loss, particularly in semi-arid agricultural regions increasingly affected by extreme rainfall. Timely crop damage assessment is critical for disaster response and insurance-related decision-making, but direct yield-loss observations are often unavailable during or shortly after flooding. This study proposes a phenology-guided regression framework for early crop damage assessment using multi-source SAR–optical observations. The study was conducted on the Tumochuan Plateau, Inner Mongolia, China, where severe rainfall beginning on 23 July 2025 caused widespread cropland inundation. Sentinel-2 EVI time series from 2022 to 2025 were fitted using a Savitzky–Golay (SG) filter, and annual area under the EVI curve (AUC) loss in 2025 relative to the 2022–2024 historical mean was used as a proxy for flood-induced crop damage. Optical features from Landsat-8/9 and Sentinel-2, together with SAR backscatter features from Sentinel-1, Lutan-1, and Gaofen-3, were incorporated into machine learning regression models. SAR features improved pixel-wise prediction, with the Random Forest model achieving the highest R2 of 0.62 using early-period features and 0.77 using later-period features. Village-scale aggregation further improved performance, yielding an early-period R2 of 0.84 across 123 and 0.78 across 122 villages. These results demonstrate the feasibility of SAR–optical and phenology-guided regression for early crop damage assessment under long-duration inundation.
Why it matches plant phenotyping methodsSAR・光学画像とフェノロジー指標を用いて作物被害を推定する回帰手法が研究の中心であり、作物状態(洪水被害)を定量化・検証しているため。
abstractThis study proposes a phenology-guided regression framework for early crop damage assessment using multi-source SAR–optical observations.
Agriculture plays a pivotal role in ensuring global food security, economic stability, and sustainable development. Plant diseases significantly reduce agricultural productivity, resulting in substantial economic losses and threatening food supply worldwide. Early and accurate identification of plant leaf diseases enables timely intervention, minimizes crop damage, and enhances agricultural yield. Traditional disease diagnosis relies heavily on visual inspection by agricultural experts, making the process labor-intensive, subjective, and unsuitable for large-scale deployment. Recent advances in artificial intelligence, particularly deep learning, have transformed plant disease diagnosis by enabling automatic feature extraction and highly accurate image-based classification. This review presents a comprehensive analysis of recent developments in deep learning techniques for plant leaf disease identification. Various convolutional neural network (CNN) architectures, including AlexNet, VGGNet, ResNet, DenseNet, EfficientNet, MobileNet, Inception, and Xception, are critically reviewed along with modern transformer-based models such as Vision Transformer (ViT), Swin Transformer, and hybrid CNN–Transformer frameworks. The paper also examines transfer learning strategies, object detection methods including YOLO and Faster R-CNN, and semantic segmentation approaches such as U-Net and DeepLabV3+. Publicly available benchmark datasets, including PlantVillage, PlantDoc, AI Challenger, Cassava Leaf Disease, and Rice Leaf Disease datasets, are discussed in terms of dataset diversity, annotation quality, and practical applicability. Furthermore, image preprocessing techniques, data augmentation methods, evaluation metrics, and deployment considerations for mobile and edge devices are comprehensively reviewed. The paper identifies current research challenges, including dataset imbalance, environmental variability, model interpretability, computational complexity, and limited real-world generalization. Finally, emerging research directions such as explainable artificial intelligence, federated learning, multimodal learning, self-supervised learning, lightweight architectures, and edge AI are discussed to provide future research opportunities. This review serves as a valuable resource for researchers, practitioners, and agricultural technologists interested in developing robust, scalable, and intelligent plant disease identification systems.
Why it matches plant phenotyping methods植物葉の病害状態を画像から識別する深層学習手法を中心に、モデル、データセット、評価、展開を体系的にレビューしており、植物表現型計測手法のレビューに該当する。
abstractThis review presents a comprehensive analysis of recent developments in deep learning techniques for plant leaf disease identification.
Early and accurate identification of plant diseases is essential for improving crop productivity and ensuring food security. Many existing deep learning-based plant disease classification methods rely solely on leaf images collected from a controlled environment, which limits their applicability in real-world agricultural conditions where symptoms may be visually unclear and influenced by environmental factors. To address these challenges, this study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification. The proposed architecture employs EfficientNet-B0 for visual feature extraction, BERT for semantic representation of symptom descriptions, and a lightweight multilayer perceptron for modeling environmental factors such as temperature, humidity, rainfall, and soil moisture. Features from all three modalities are fused into a unified representation to train the CNN model. The model is trained and tested upon the Context-Aware Multimodal Augmented PlantVillage dataset covering 38 plant diseases and healthy classes. Experimental results show that AgriFusionNet gives an overall accuracy of 98.94% on the dataset Context-Aware Multimodal Augmented PlantVillage, with competitive precision and recall and F1-score. The multimodal framework facilitates the co-learning of visual, semantic, and contextual environmental representations and the analyses of the confusion matrix and feature interactions give insights into cross-modal relationships. The proposed approach aims to explore context-aware multimodal representation learning for agricultural AI applications, with emphasis on integrating complementary visual, semantic, and contextual information.
Why it matches plant phenotyping methods葉画像を中心に、症状記述と環境情報を統合して植物病害状態を分類する手法を開発・評価しており、植物フェノタイピング手法が中心である。
abstractthis study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification.
Reproduction assets foundThe paper's data availability statement points to the Context-Aware Multimodal Augmented PlantVillage dataset (leaf images, symptom text, environmental data used for the phenotyping/classification analysis) deposited publicly on IEEE Dataport with a DOI matching an allowed URL.Dataset · publicPublicly available datasets were analyzed in this study. This data can be found here: Dataset. IEEE Dataport. https://dx.doi.org/10.21227/9jat-r836 [Accessed on August 2025].Open asset ↗IEEE Dataport · 10.21227/9jat-r836lines:1029-1047Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
Introduction Climate extremes increasingly threaten agricultural production, yet many artificial intelligence systems in agriculture remain local, reactive and narrowly trained for one crop, region or sensing modality. Methods We present AgriFM, a multimodal geospatial foundation model that combines satellite image time series, radar, thermal observations, weather trajectories, soil properties, topography and sparse management variables to estimate crop-stress and yield-failure risk across crops and regions. AgriFM was pretrained using self-supervised objectives on 2.4 million field-season sequences and evaluated on a curated benchmark spanning maize, wheat, soybean, rice and sorghum across five agroclimatic regions. Results In held-out geography and time-split evaluations, AgriFM improved early stress detection and yield-failure prediction over statistical, crop-model and deep-learning baselines. The largest gains occurred during compound drought and heat events, for which AgriFM produced alerts 18 to 24 days earlier than the satellite-only baseline while maintaining improved calibration. Phenology-conditioned fusion improved transfer across planting calendars, and uncertainty calibration reduced false alerts at fixed recall. Discussion Because the study is based on retrospective datasets, these findings establish cross-region retrospective performance rather than prospective field efficacy. The results support further field-based evaluation of multimodal foundation models for climate-resilient crop monitoring.
Why it matches plant phenotyping methods作物ストレス状態と収量失敗リスクを衛星・レーダー・熱画像等から推定する基盤モデルを開発し、複数作物・地域のベンチマークで評価しており、植物状態の取得・推定手法が中心である。
abstractWe present AgriFM, a multimodal geospatial foundation model that combines satellite image time series, radar, thermal observations, weather trajectories, soil properties, topography and sparse management variables to estimate crop-stress and yield-failure risk across crops and regions.
Abstract Outbreaks of plant diseases are major threats to world food security particularly in areas where real-time monitoring and quick decision support are constrained by low-power edge gadgets and untrustworthy connectivity. In order to overcome these issues, this paper presents a FPGA-Accelerated IoT implementation of a Causal-Attention Multi-Modal Deep Learning Network, named EpiFusionNet-Edge, that can be applied to monitor crop diseases with real-world farming scenarios with high precision and scalability. The framework incorporates five data modalities that are complementary in nature and they include RGB leaf pictures, microscopic foldscope images, UAV hyperspectral signatures, microclimate IoT sensor measurements and region-specific pathogen/pest pressure indexes giving a complete picture of the health of the plant. Dual causal-attention mechanism is proposed to simulate both spatial and temporal environmental factor activation, which helps to detect and make predictions at the early stage and provides an explanatory logic behind the decisions. Multi-task learning enables classification of diseases, quantification of their intensity at the level of a micro-prediction and prediction of outbreaks in the short term (1–30 days). In order to achieve deployability in resource-constrained settings, the proposed deep learning architecture is ensemble-distilled, structurally pruned, and INT8-quantized, and hardened on a Xilinx Zynq-7000 FPGA platform. The FPGA accelerator is 43.2x faster inference, 88 percent less power usage, and less than 10 ms latency, which allows real-time execution of continuous field monitoring with IoT sensors. Cross-condition assessment on multi-domain datasets shows that there are great improvements on cross-environment generalization rates with 98.6% classification accuracy, 92.7% severity estimation accuracy and less than 3.5% degradation with domain shift. Grad-CAM + + and causal feature traceability further add interpretability with the focus of the model and the pathological indicators proven by experts. The findings show the promise of using a combination of IoT sensing, multi-modal AI fusion, and FPGA hardware acceleration to develop a deployable and scalable and transparent system with regard to precision agriculture. This paper creates a roadmap to a new generation of smart farming systems that are able to conduct disease surveillance and actively protect crops at the periphery in an autonomous manner.
Why it matches plant phenotyping methods植物病害の画像・センサー観測から病害強度を定量化するマルチモーダル・エッジ推論基盤を開発し、精度・速度・消費電力・ドメインシフトを評価しているため、植物フェノタイピング手法が中心である。
abstractthis paper presents a FPGA-Accelerated IoT implementation of a Causal-Attention Multi-Modal Deep Learning Network, named EpiFusionNet-Edge
TomatoMultimodalLeafClassificationDisease symptoms / severityStress response / tolerancePlant / canopy temperature
Wearable plant sensing systems for simultaneous biochemical and physiological monitoring with real-time multimodal data analysis remain limited. Here, we present FolioClip, a multimodal wearable patch that continuously monitors leaf temperature, humidity, light, CO 2 , and three volatile organic compounds (VOCs) with high selectivity. Its bookmark-inspired design enables secure attachment to plant leaves of diverse morphologies and it integrates a flexible printed circuit board for wireless data transmission. We also develop FolioOmni, an open-source machine learning (ML) framework for sensor importance ranking, multi-stress classification, and early stress detection. The integrated FolioClip–FolioOmni platform detects and classifies nine stresses, including light, water, CO 2 , mechanical cut, P. infestans , and A. alternata , in tomato plants with 92% accuracy. Notably, P. infestans on tomato was detected within 15.5 h post-inoculation, earlier than quantitative polymerase chain reaction (qPCR) (~4 days) and visual phenotyping (~7 days), highlighting the potential of integrating multimodal wearable sensing and online ML for precision agriculture.
Why it matches plant phenotyping methods葉の生理・健康状態とストレスを測定するウェアラブルセンシング装置および機械学習解析基盤を開発し、複数ストレスで性能評価しているため、植物フェノタイピング手法が中心である。
abstractWe also develop FolioOmni, an open-source machine learning (ML) framework for sensor importance ranking, multi-stress classification, and early stress detection.
Background Salinity is a major abiotic stress that negatively affects nearly all plant species at all stages of growth. Drought and poor-quality irrigation cause high soil salinity and salt accumulation via evaporation, reducing crop productivity. Despite its critical importance, the spatial localization of salt ions and associated biochemical changes within plants experiencing high salinity remains largely unknown. In this study, we developed a multimodal imaging pipeline to understand the impact of salinity on the pistachio rootstock UCB-1 (Pistacia atlantica x Pistacia integerrima). We directly link biochemical fingerprints in stem tissue architecture with salt ion localization to provide insights into the strategies pistachio uses to tolerate salinity. Results We observed that Pistacia spp. exposed to high salt conditions accumulated Ca, Si, Cl, Al and Mg as hotspots within the pith, compared to the control (of which only Ca and Al co-locate). In contrast, there was a decrease in K between the control and salinity treatment. Hotspots of amide I and II were present in the cortex and pith of the salinity treated sample. Additionally, the salinity treatment resulted in an increased abundance of pectin and carbohydrates within the pith compared to the control, and the abundance of esters/carboxylic acid was greater in the salinity treatment. Conclusions We determined that Cl and K, S and P, and biochemical components polysaccharide and pectin, esters and carboxylic acid, amide I and cellulose are the strongest drivers of salinity-treatment induced variability. In the cortex and phloem/xylem, a negative K-Ca correlation decreases in the salinity treatment. Several hotspots of elements and amide I (proteins) appear under salinity treatment, particularly in the cortex, suggesting an increase in the production of stress-related proteins (in response to high Cl) and/or structural proteins (i.e. Ca). Together, these results indicate that pistachio responds to salinity through ion compartmentalization coupled with a targeted biochemical adjustment, rather than a broadscale tissue-wide response. Overall, these novel, spatially resolved pixel-registered multimodal imaging data provide an enabling platform to understand the mechanisms of salinity tolerance in Pistacia spp and can be broadly applied to studying stress-related phenotype response in various plant tissues.
Why it matches plant phenotyping methods植物組織の元素・生化学状態を空間的に取得するピクセル登録型マルチモーダル画像パイプラインを開発し、植物ストレス表現型解析への汎用的プラットフォームとして提示しているため。
abstractwe developed a multimodal imaging pipeline to understand the impact of salinity on the pistachio rootstock UCB-1
Accurate crop stress detection is essential for precision agriculture; however, most existing approaches rely on binary labels that collapse distinct stress processes water deficit, nutrient deficiency, disease, and pest damage into a single "stressed" category.We demonstrate empirically that this binary formulation is the primary barrier to classification performance: five model architectures achieve ROC-AUC values within ±0.01 of the random baseline (0.50) on binary stress classification, regardless of feature engineering strategy. Decomposing the binary label into stress-type-specific categories enables anXGBoost classifier to achieve 91.4% accuracy and a macro-averaged F1-score of 0.93 using the same underlying features.To extend coverage to visual disease symptoms, we train a MobileNetV2-based CNN on paddy leaf images, achieving 93.7% binary accuracy (healthy vs. disease_stress) with 100% healthy recall.We combine both modalities in a fusion ensemble that merges tabular and image predictions through rule-based priority logic, achieving 94.6% accuracy on the evaluated image subset.
Why it matches plant phenotyping methods葉画像から健全・病害ストレス状態を推定するCNNと、画像・表形式データの融合分類法が研究の中心であり、植物状態の取得・推定手法を評価している。
abstractTo extend coverage to visual disease symptoms, we train a MobileNetV2-based CNN on paddy leaf images, achieving 93.7% binary accuracy (healthy vs. disease_stress) with 100% healthy recall.
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
CottonField / plotMultimodalFruitClassificationPhysiological trait estimationGrowth / development / phenology
Cotton fiber quality is shaped during boll development, boll opening, fluffing, and harvesting, but current assessment still relies largely on manual field inspection and postharvest laboratory testing. This limits timely harvest scheduling and plot-level quality management. To address this problem, we propose a self-supervised multimodal sensing framework for linking preharvest cotton boll status, environmental conditions, and postharvest fiber quality. First, the Cotton Boll Visual Phenotype Self-Supervised Encoding Module learns maturity-related visual representations by reconstructing masked image patches, so that boll cracking, lint exposure, and surface texture can be captured from unlabeled field images. Second, the Agricultural Sensor Temporal Masked Modeling Module reconstructs masked sensor observations to model temporal patterns in temperature, humidity, light, soil moisture, rainfall, and other environmental variables. Third, the Vision–Environment Cross-Modal Contrastive Fusion Module aligns image features with environmental features and produces a joint representation for downstream prediction. Field experiments were conducted using cotton boll images from different maturity and abnormal states, environmental sensor records, management information, and postharvest fiber quality measurements. The framework was evaluated for maturity classification, harvest-window recognition, and fiber quality prediction. The results showed that the proposed method performed consistently better than representative machine learning, single-modal deep learning, and multimodal fusion baselines, while few-shot and ablation experiments supported the value of self-supervised pretraining and multimodal fusion. These findings indicate that the proposed approach can provide useful information for preharvest cotton maturity assessment and harvest-quality management.
Why it matches plant phenotyping methods綿花の成熟状態を画像・環境センサーから抽出し、成熟度分類や収穫時期認識を行うマルチモーダル手法の開発・評価が中心であり、植物状態の表現型推定に該当する。
abstractwe propose a self-supervised multimodal sensing framework for linking preharvest cotton boll status, environmental conditions, and postharvest fiber quality.
India’s economy is primarily based on agriculture. Agriculture has significant contribution in nation’s GDP. Food security and employment significantly influenced by agriculture. However factors like uncertain weather conditions, poor quality of seeds and plant diseases impact on agriculture productivity. Computer vision and DL algorithms are most crucial components of precision agriculture. Early detection can improve decision making, maximize pesticide use, and preserve harvests. Using CNN architectures, segmentation-based approaches, handcrafted feature-based methods, and hybrid approaches incorporating Machine Learning and Deep Learning this study seek to provide review of recent publications from 2020 to 2026. The review was carried out using a variety of publications with different datasets, methodologies, and outcomes. The findings show that DL, especially CNN and transfer learning models, performed better than machine learning techniques. It points out several significant problems, such as dataset imbalance, insufficient generalization, computing inefficiency, and a dearth of real-world data. Future research topics are also suggested which includes IoT-driven real-time solutions, lightweight architecture, domain adaption, and multimodal imaging. This review aims to develop plant disease detection technologies that are more dependable, scalable, and field deployable.
Why it matches plant phenotyping methods植物葉の病徴を画像から検出する機械学習・深層学習手法をレビューおよび実験的に扱っており、植物フェノタイピング手法が中心である。
titlePlant Leaf Disease Detection Using Machine Learning and Deep Learning: A Review and Experimental Study
Artificial intelligence (AI) has emerged as a transformative tool for plant health monitoring, offering new opportunities for scalable, timely, and data-driven pest and disease management in agriculture. This review provides a comprehensive synthesis of AI-based methods for pest and plant disease detection, systematically organizing existing literature across sensing modalities, learning paradigms, and deployment scales. We distinguish between population-level pest monitoring, plant-centric visual inspection, and field-scale surveillance, as well as between post-symptomatic disease recognition and pre-symptomatic detection enabled by spectral imaging technologies. Beyond summarizing recent advances, this work places strong emphasis on critical analysis, discussing fundamental limitations related to data scarcity, domain shift, generalization under field conditions, and the challenge of disentangling biotic from abiotic stress factors. The review further examines the distinction between correlation-driven AI predictions and causal disease understanding, positioning AI as a complementary decision-support tool alongside established diagnostic methods. Building on these insights, we outline key future research directions, including multimodal sensor fusion, explainable and trustworthy AI, edge-based deployment for real-time monitoring, and the development of foundation models for unified agricultural intelligence. This review aims to serve as both an accessible entry point and a critical reference for advancing AI-driven plant health management.
Why it matches plant phenotyping methods植物の病害・害虫状態を画像・スペクトルなどで検出するAI手法を対象としたレビューであり、植物状態の取得・推定手法が中心である。
titleA Review on Artificial Intelligence Methods for Plant Disease and Pest Detection
This research introduces a multimodal deep learning framework for early detection of plant pathogens to capture pre-symptomatic biochemical changes in plants while simultaneously modeling the environmental drivers of disease development. A hybrid fusion architecture combines 3D convolutional neural networks for spatial-spectral feature extraction from HSI cubes with transformer-primarily based temporal modeling of climate sequences. Cross-modal attention mechanisms dynamically weight discriminative features, which includes chlorophyll degradation bands and humidity thresholds, to permit joint representation learning. The framework achieved 94.5% accuracy in pathogen detection, outperforming unimodal HSI (84.1%) and climate- only (76.5%) baselines by 10-18 percentage points. Moreover, it detected fungal infections 5-7 days before visual symptom onset and had a 12.3% higher F1-rating compared to the current methods. Field simulations showed that precision application resulted in 41% reduction in fungicide use. By connecting proximal sensing with climatic analytics, this research contributes to precision agriculture by providing timely and eco-friendly pest control of diseases. The multimodal fusion framework is introduced to overcome the limitations of unimodal approaches. It integrates the most appropriate data sources, thus allowing the earliest and most accurate detection of plant pathogens.
Why it matches plant phenotyping methods植物の病害状態をハイパースペクトル画像から抽出するマルチモーダル手法の開発・評価が中心であり、単なる病原体診断や農薬施用試験ではない。
abstractThis research introduces a multimodal deep learning framework for early detection of plant pathogens to capture pre-symptomatic biochemical changes in plants
High-throughput plant phenotyping generates valuable data that often remains trapped in unstructured text and isolated RGB images. To bridge this semantic gap, we propose a framework for constructing a multimodal granular Knowledge Graph (KG) to monitor genotype-phenotype interactions across time and experiments. In this work, we focus on wheat Triticum aestivum as a representative target crop to validate our methodology across complex canopy environments. Our pipeline first distills noisy field notes to extract entities and relations, dynamically constructing the KG by converting unique instances into hierarchical class entities via RDF-typing. These graph nodes are then aligned with standardized ontologies (PO, RO, WTO) using PlantDeBERTa. To visually ground the constructed graph, a Vision-Language Model paired with a wheat-segmentation ViT generates attention-based softmaps, linking specific KG entities directly to image pixels. We introduce a central observation node Plant_Obs_Id to connect these multimodal subgraphs temporally. Evaluated on 500 curated WisWheat samples using Pointing Game accuracy, Visual Word Sense Disambiguation (VWSD), and rank-based metrics, our neuro-symbolic approach successfully maps complex field observations to a structured graph. This enables automated field note auditing, temporal stress monitoring, and precise spatial trait localization for wheat breeders.
Why it matches plant phenotyping methods植物のマルチモーダル表現型データを知識グラフと画像に統合し、画像画素への形質局在化を行う中核的な計算フレームワークを提案・評価しているため。
abstractwe propose a framework for constructing a multimodal granular Knowledge Graph (KG) to monitor genotype-phenotype interactions across time and experiments
Plant diseases remain a major challenge to global food production, and timely, accurate, and scalable detection of plant stress is critical to reducing these losses. Recent advances in digital imaging and artificial intelligence offer unprecedented opportunities for precision crop disease detection and management. Yet, existing plant disease datasets remain often fragmented across crop and disease systems, and are largely dominated by controlled-environment imagery. The lack of standardized, interoperable, and representative datasets limits reproducibility, transferability, and scalability of AI systems, thereby constraining their deployment in operational agricultural applications. Here we present LeafMD, an integrated multimodal plant disease dataset and benchmark resource that includes LeafNet 2.0, a large-scale multimodal digital image dataset comprising 255,855 image–text pairs across 37 crop species, 197 crop–disease classes, and 9 geographic regions spanning tropical, subtropical, and temperate agricultural systems. Unlike conventional datasets, LeafNet 2.0 integrates biologically grounded symptom descriptions with image-level annotations of early and late disease stages, enabling symptom-aware analysis of disease progression under realistic field conditions. We further introduce LeafBench 2.0 as part of LeafMD, a visual-question answering benchmark covering nine fine-grained plant pathology tasks, including pathogen classification, lesion characterization, symptom interpretation, and disease severity assessment. Evaluation across 16 vision–language models revealed substantial performance gaps between coarse disease recognition and fine-grained pathological reasoning, while agriculture-adapted models consistently outperformed several larger general-domain architectures on symptom-oriented tasks. Together, LeafNet 2.0 and LeafBench 2.0 establish LeafMD as a multimodal resource for developing disease-aware agricultural foundation models and studying fine-grained pathological reasoning in real-world environments.
Why it matches plant phenotyping methods植物病害の画像・症状記述データセットとベンチマークを構築し、病徴解釈・病変特徴・病害重症度評価を対象にモデル性能を評価しており、植物状態の取得・評価手法が中心である。
abstractHere we present LeafMD, an integrated multimodal plant disease dataset and benchmark resource
Accurate perception and 3D reconstruction of fruit tree branch structures are fundamental to smart orchard development, with broad applications in intelligent harvesting, crop phenotyping, and precision management. However, the slender and highly branched morphology, multi-scale distribution, weak surface texture, and severe occlusion inherent to fruit tree branches pose substantial challenges to high-fidelity modeling. This paper systematically reviews advances in branch feature extraction and 3D reconstruction for fruit tree canopies. A structured literature search was conducted using the Web of Science, Scopus, and Google Scholar databases, with search terms including “fruit tree branch”, “point cloud reconstruction”, “3D canopy modeling”, “branch feature extraction”, and “agricultural robotics”. Studies published between 2000 and 2025 were considered, with inclusion criteria requiring relevance to branch structure perception, reconstruction accuracy, or orchard application; non-peer-reviewed sources and studies lacking quantitative evaluation were excluded. We trace the evolution of feature extraction from classical 2D image processing and geometric fitting, through point cloud segmentation and skeleton extraction, to modern deep learning approaches and multimodal perception techniques. For 3D reconstruction, we compare active and passive sensing strategies alongside both explicit and implicit scene representation methods, discussing their respective strengths and applicable scenarios. A five-dimensional evaluation framework is also proposed, encompassing geometric accuracy, structural consistency, feature stability, computational efficiency, and generalization capability. Finally, we identify key bottlenecks in fine-grained structure recovery, occlusion handling, and cross-scene generalization, and highlight future directions in structural prior integration, multimodal collaborative modeling, and lightweight neural representations—offering a structured reference for advancing 3D perception research in smart orchards.
Why it matches plant phenotyping methods果樹の枝構造の特徴抽出と3D再構成を対象とする、植物形態計測・表現型取得手法のレビューであり、方法論が中心です。
abstractThis paper systematically reviews advances in branch feature extraction and 3D reconstruction for fruit tree canopies.
Field / plotMultimodalLiDAR / point cloudMultispectral / hyperspectralWhole plant / canopy / plot / fieldArchitecture / morphology / geometry
Crop phenotyping serves as a fundamental basis for crop breeding, precision cultivation, and smart agriculture. In recent years, it has evolved toward multi-modal integration and multi-scale coordination. This paper analysed indoor and outdoor phenotyping platforms across diverse application scenarios, and reviewed sensing technologies including RGB imaging, multi-spectral imaging, hyperspectral imaging, thermal imaging, fluorescence imaging, LiDAR, and nuclear magnetic resonance (NMR). The applications of these technologies were summarized in capturing crop morphological traits, physiological status and biochemical components. The phenotyping acquisition methods and intelligent analytical techniques were also analyzed at different scales such as plant cells, tissues and organs, individual plants, population plot and field. Additionally, the advancements were explored in high-throughput phenotyping technologies and their integration with crop gene function analysis, providing a reference for future phenotyping research.
Why it matches plant phenotyping methods作物表現型センシング技術・装置、取得法、解析技術、プラットフォームを主題とする包括的レビューであり、表現型手法が中心です。
abstractThis paper analysed indoor and outdoor phenotyping platforms across diverse application scenarios, and reviewed sensing technologies including RGB imaging, multi-spectral imaging, hyperspectral imaging, thermal imaging, fluorescence imaging, LiDAR, and nuclear magnetic resonance (NMR).
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
Agricultural production in arid regions is strongly constrained by water stress, making timely evaluation of crop water conditions increasingly important. However, conventional measurements of plant moisture content (PMC) primarily rely on destructive oven-drying methods, which are not only labor-intensive and time-consuming but also constrained by limited sample size and spatial coverage. These shortcomings make it difficult to capture the spatial heterogeneity of crop water status across large agricultural regions, thereby restricting regional-scale water diagnosis and precision irrigation decision-making. Focusing on silage maize cultivated in the arid region of Gansu Province, China, this work develops a regional PMC estimation approach by combining multi-source remote sensing data. High-resolution unmanned aerial vehicle (UAV) observations were integrated with Sentinel-2 and Sentinel-3 imagery, while radiometric and temperature corrections were applied to improve data consistency. A set of spectral, textural, and thermal features was derived from multispectral, visible, and thermal infrared datasets. Feature selection based on Pearson correlation was then carried out, followed by the construction of three models, namely Random Forest (RF), Support Vector Machine (SVM), and Partial Least Squares Regression (PLSR). Among them, the RF model performed more reliably, achieving a validation R2 of 0.92 with relatively low prediction error. In addition, calibration using UAV data led to a clear improvement in satellite-based estimates, with R2 increasing from 0.52–0.62 to 0.71–0.74. The generated PMC maps captured both the temporal decline during the growing season and the spatial variability across the study area. Overall, the proposed approach offers a practical option for large-scale monitoring of crop water status and can support irrigation management in water-limited environments.
Why it matches plant phenotyping methodsマルチソースリモートセンシングと機械学習により、トウモロコシの植物含水量という明示的な植物状態を地域スケールで推定・検証する手法開発が中心である。
abstractthis work develops a regional PMC estimation approach by combining multi-source remote sensing data.
Abstract has not been obtained from indexed metadata or an accessible article page.
Why it matches plant phenotyping methods植物フェノタイピングのためのマルチモーダル融合・深層学習フレームワーク開発が題名上の中心であり、土壌推定も併記されているが、植物表現型抽出手法を含むため採用。
titleUnified Hybrid Deep Learning Framework for Multi-Modal Data Fusion in Plant Phenotyping and Soil Nutrient–Moisture Prediction using BiLSTM–Transformer Networks
Early and accurate plant disease diagnosis is essential for sustainable crop production and effective disease management. This study presents an Explainable Gradient-Based Convolutional Neural Network (EG-CNN) that integrates leaf image data with gene expression and metabolomics information to enhance disease classification while providing transparent and interpretable predictions. The proposed framework was evaluated on four major plant diseases, powdery mildew, blight, anthracnose, and leaf spot, using a multimodal dataset comprising 10,000 images and associated omics data. Comparative analysis against Traditional CNN, ResNet-50, Vision Transformer (ViT), and CNN-LSTM models demonstrated the superior performance of EG-CNN. The model achieved 97.4% accuracy, 97.1% precision, 96.8% recall, and a 96.9% F1-score, outperforming all benchmark approaches. Paired t-test results revealed statistically significant improvements (p 0.05) over competing models. Confusion matrix analysis indicated minimal misclassification across disease classes, while ROC analysis produced near-perfect AUC values, confirming excellent class separability, sensitivity, and specificity. Five-fold cross-validation further demonstrated robust generalization, with accuracy ranging from 97.1% to 97.6%. Grad-CAM visualizations successfully identified disease-relevant regions, including lesions, discoloration, and necrotic tissues, thereby enhancing model transparency and user trust. The integration of omics data also improved biological interpretability by linking predictions to underlying molecular responses. In conclusion, the proposed EG-CNN framework provides a highly accurate, robust, and explainable solution for plant disease diagnosis. Its multimodal architecture and strong generalization capability make it a promising tool for precision agriculture and real-time disease monitoring applications.
Why it matches plant phenotyping methods植物葉画像から病害症状を推定する説明可能な深層学習フレームワークを開発し、複数モデルとの比較、交差検証、Grad-CAMによる病徴領域の評価を行っており、植物表現型取得・判定手法が中心である。
abstractThis study presents an Explainable Gradient-Based Convolutional Neural Network (EG-CNN) that integrates leaf image data with gene expression and metabolomics information to enhance disease classification while providing transparent and interpretable predictions.
Accurate yield estimation and crop load monitoring are essential for precision orchard management, supporting targeted fertilization, pruning, thinning, harvest planning, and marketing decisions. However, reliable in-situ monitoring remains challenging because commercial orchards are characterized by severe canopy occlusion, fruit overlap, heterogeneous tree architecture, variable illumination, and complex backgrounds. This review synthesizes advances in multi-modal sensing and deep learning for orchard yield estimation, breaking down the paradigm into intermediate fruit-counting or crop-load monitoring steps and supplementary spectral quality-assessment dimensions. First, yield-related indicators are summarized, including direct phenotypic traits such as fruit number, size, volume, and spatial distribution, as well as indirect structural and physiological proxies such as canopy volume, vegetation indices, flowering intensity, and spectral maturity attributes. Second, representative sensing devices and carrying platforms are reviewed, including red-green-blue (RGB) cameras, red-green-blue-depth (RGB-D) sensors, light detection and ranging (LiDAR), hyperspectral and multispectral systems, unmanned ground vehicles (UGVs), and unmanned aerial vehicles (UAVs). Third, the evolution of estimation methods is discussed, from traditional image processing and machine learning to object detection, instance segmentation, multi-object tracking, point-cloud analysis, remote-sensing regression, and multi-modal fusion. The review shows that no single sensor or algorithm can satisfy all orchard monitoring requirements. Ground-based vision and depth sensing are more suitable for fine-scale fruit counting and sizing, whereas UAV and spectral sensing provide advantages for regional yield mapping and quality-enhanced assessment. Future research should emphasize occlusion-aware perception, robust cross-environment generalization, lightweight edge deployment, standardized benchmarks, and integrated quantity-quality monitoring frameworks for actionable crop load management.
Why it matches plant phenotyping methods果実数・サイズ・体積・空間分布などの植物形質を対象に、センシング機器と画像解析・深層学習による収量推定法を体系的にレビューしており、フェノタイピング手法が中心である。
abstractThis review synthesizes advances in multi-modal sensing and deep learning for orchard yield estimation, breaking down the paradigm into intermediate fruit-counting or crop-load monitoring steps and supplementary spectral quality-assessment dimensions.
Tapioca is a vital root crop whose yield is influenced by multiple environmental, soil, and physiological factors.Traditional yield estimation methods depend upon statistical models and manual sampling, which often lack real-time accuracy and are time-consuming and labour-intensive.The combination of Internet of Things (IoT) sensor networks and Unmanned Aerial Vehicle (UAV)/satellite imagery, with deep learning (DL) models, gives a promising solution for accurate real-time yield prediction.In this study, sensor data (soil moisture, weather parameters, pH, NPK levels, temperature, and leaf chlorophyll content) are collected using devices like Davis Vantage Pro2, Decagon 5TE and SPAD-502, while image data are captured using DJI Phantom 4 Multispectral UAVs and Sentinel-2 satellite imagery.Pre-processing of sensor data involves missing value imputation, normalization, and feature selection using the Hybrid Frilled Lizard Osprey (HFLO) algorithm, while image data undergo resizing, augmentation, and filtering approach.Image based features are extracted by a position-attention DenseNet-201 model.Also, the features from image data and sensor data are fused using a concatenation mechanism.Finally, fully connected layers and improved support vector regression (ISVR) are used to identify the tapioca yield prediction.The integrated DL methods give higher accuracy than individual modalities for both modalities.The image-based model captures spatial and phenotypic variations, while the sensor-based model captures fine-grained environmental effects.The proposed approach obtained the MAE value of 0.0566, RMSE value of 0.654, and R 2 value of 99.5, demonstrating the potential of multimodal real-time data integration for precision agriculture in tapioca fields.
Why it matches plant phenotyping methodsUAV・衛星画像と環境センサーを統合し、深層学習でキャッサバの収量という植物形質を推定する手法の開発が中心である。
abstractThe combination of Internet of Things (IoT) sensor networks and Unmanned Aerial Vehicle (UAV)/satellite imagery, with deep learning (DL) models, gives a promising solution for accurate real-time yield prediction.
Against the backdrop of the rapid development of smart agriculture, pest and disease monitoring and crop growth assessment for large-scale farmlands are of substantial importance for precision management and risk early warning. However, traditional unimodal visual methods are highly susceptible to illumination variation, canopy occlusion, scale differences, and background interference in real field environments, and thus fail to make full use of environmental sensing information and spatial priors. To address these issues, a multimodal target perception framework for intelligent farmland inspection is proposed in this study. By jointly integrating UAV imagery, time-series data from ground Internet of Things sensors, and spatial positional information, joint modeling of pest and disease recognition and crop growth assessment is achieved through cross-modal alignment and collaborative encoding, multi-scale target perception, and dynamic multimodal fusion and decision-making. Experimental results demonstrate that, in the pest and disease recognition task, the proposed method achieved a Precision of 91.63%, a Recall of 90.27%, an F1-score of 90.94%, and an mAP of 93.15%, significantly outperforming comparison models such as Faster R-CNN with ResNet50 backbone, YOLOv8-m, Swin Transformer-Tiny, and Multimodal Transformer. In the crop growth assessment task, an Accuracy of 89.96%, a Precision of 89.11%, a Recall of 88.74%, and a Macro-F1 of 88.92% were achieved, again clearly exceeding those of ResNet50, EfficientNet-B3, ViT-B/16, and conventional multimodal fusion models. The ablation study further verified the effectiveness of the cross-modal alignment module, the multi-scale target perception module, and the dynamic fusion module, with the complete model reaching 90.94%, 93.15%, and 88.92% in Pest F1, Pest mAP, and Growth Macro-F1, respectively. Furthermore, the net economic return regression experiment at the unit-area level further demonstrates that the proposed method can effectively connect state information with economic outcomes, showing strong application potential in return prediction, performance evaluation, and resource allocation optimization. These findings indicate that the proposed method can effectively improve perception accuracy and robustness in complex farmland environments, thereby providing reliable technical support for intelligent inspection, pest and disease early warning, and precision management in agricultural scenarios.
Why it matches plant phenotyping methodsUAV画像、IoT時系列データ、空間情報を統合したマルチモーダル手法を開発し、作物生育状態の評価を技術的に検証している。害虫認識単独ではなく、植物の生育評価を含む取得・推定手法が中心である。
abstracta multimodal target perception framework for intelligent farmland inspection is proposed in this study.
In modern agriculture, artificial intelligence (AI) is doing excellent work in crop monitoring, crop disease detection, crop yield prediction, and crop stress assessment. Various techniques such as deep learning, generative models, vision transformers, explainable AI (XAI), multimodal fusion, etc., have helped in building intelligent crop analysis. This paper provides the comparative crop analysis of various crop species, data modes, and environmental conditions for the review of benchmark studies for the current framework, experimental methodologies, and datasets. The important challenges and open issues identified are limited field datasets, class imbalance, dataset bias, high computational complexity, privacy concerns, etc. Based on these, we suggested future work that can include foundation models, digital twin techniques, federated learning, multimodal frameworks, and interpretability architecture. This review provides a review for creating reliable, scalable, and sustainable AI-driven crop analysis systems. In addition to that, the survey seeks to give researchers and AI practitioners a comprehensive analysis of the current situation.
Why it matches plant phenotyping methods作物の病害・収量・ストレス評価を対象に、AI手法、データモード、ベンチマーク、実験方法、データセットを体系的にレビューしており、植物表現型取得・推定手法のレビューが中心です。
abstractThis paper provides the comparative crop analysis of various crop species, data modes, and environmental conditions for the review of benchmark studies for the current framework, experimental methodologies, and datasets.
Introduction eaf area index (LAI) and leaf nitrogen accumulation (LNA) are key indicators of wheat growth and nitrogen nutritional status. However, existing prediction methods predominantly rely on single-modal information and single-output models, limiting their ability to characterize the complex structural and physiological traits of crops. This study aimed to develop a multimodal learning framework for the simultaneous and accurate prediction of wheat LAI and LNA. Methods Spectral, image, and canopy structural features were extracted from wheat canopies across different cultivars, nitrogen treatments, and growth stages. A canopy height correction-based preprocessing method was developed to improve the extraction of structural features. A Dual-Output Bayesian Neural Network (DO-BNN) was then constructed to simultaneously predict LAI and LNA. In addition, an Extreme Sample Mining (ESM) strategy and a joint loss function were introduced to strengthen the learning of complementary information across modalities and the intrinsic correlation between the two target variables. Results The DO-BNN achieved its best predictive performance when all feature modalities were fused. The coefficients of determination (R²) for LAI and LNA were 0.89 and 0.77, respectively, while the corresponding relative root mean square errors (RRMSEs) were 0.15 and 0.35. Compared with single-modal and conventional single-output approaches, the proposed method provided more accurate and robust predictions of both wheat growth parameters. Discussion The results demonstrate that integrating spectral, image, and structural information can improve the characterization of wheat canopy traits. By jointly modeling LAI and LNA, the DO-BNN effectively exploited the physiological relationship between crop growth and nitrogen accumulation. The proposed framework provides a promising approach for the high-accuracy, collaborative monitoring of wheat growth and nitrogen nutritional status.
Why it matches plant phenotyping methods小麦キャノピーのスペクトル・画像・構造情報からLAIと葉窒素蓄積を推定するマルチモーダル手法を開発し、前処理、ニューラルネットワーク、性能比較まで中心的に扱っているため。
abstractThis study aimed to develop a multimodal learning framework for the simultaneous and accurate prediction of wheat LAI and LNA.
TomatoMultimodalLeafClassificationStress / disease detectionStress response / tolerancePlant / canopy temperature
Wearable plant sensing systems for simultaneous biochemical and physical monitoring with real-time multimodal data analysis remain limited. Here, we present PhytoClip, a multimodal wearable patch that continuously monitors leaf temperature, humidity, three volatile organic compounds (VOCs) with high selectivity, and microenvironmental light intensity and CO2 concentration. PhytoClip features a bookmark-inspired design for secure attachment to leaves of diverse morphologies, supported by a flexible printed circuit board for data acquisition, wireless communication, and cloud-based monitoring. We develop PhytoSense, an open-source machine learning (ML) framework for sensor importance ranking, multi-stress classification, and early stress detection. The integrated PhytoClip-PhytoSense platform detects and classifies nine biotic and abiotic stresses in tomato plants with 92% accuracy. Notably, P. infestans on tomato was detected within 15.5 h post-inoculation, earlier than quantitative polymerase chain reaction (qPCR) (~4 days) and visual phenotyping (~7 days), highlighting the potential of integrating multimodal wearable sensing and online ML for precision agriculture.
Why it matches plant phenotyping methods植物の葉に装着するマルチモーダルセンサーとオンラインMLによるストレス・病害状態の取得および分類が研究の中心であり、植物フェノタイピング手法として明確に該当する。
abstractWe develop PhytoSense, an open-source machine learning (ML) framework for sensor importance ranking, multi-stress classification, and early stress detection.
A comprehensive intelligent system for corn yield prediction based on the synergy of a Spatio-Temporal Generative Adversarial Network (ST-cGAN) and a recurrent CNN-LSTM architecture has been developed and tested. A multimodal fusion of weather-independent Sentinel-1 radar data, ERA5-Land meteorological factors, and historical Sentinel-2 optical observations was applied. The problem of "biochemical blindness" in radar signals, where radar captures physical plant structure but fails to detect photosynthetic activity, and cloud cover limitations in optical remote sensing was successfully resolved. To achieve this, an early fusion strategy was implemented. The ST-cGAN simultaneously processes current SAR data for the structural macro-state, historical optical data for the last known biochemical baseline, and a 10-day history of meteorological priors to enforce physiological constraints. Consequently, if radar indicates high biomass but meteorological data reveals severe drought, the neural network mathematically recognizes biological stress and synthesizes proportionally depressed NDVI and NDRE indices, preventing the hallucination of falsely healthy crops. Dynamic synthesis of missing NDVI and NDRE vegetation indices was conducted under prolonged continuous cloud cover (up to 21 days), achieving a high structural similarity index (SSIM = 0.89). Temporal discriminators for frame sequence analysis were introduced into the architecture, improving the edge-preserving index by 18 % and minimizing spatial artifacts at field boundaries. Pixel-level regression was performed for a 15,000-hectare test area in the Western Forest-Steppe of Ukraine based on reconstructed time series covering 15 critical phenological stages. It was established that the proposed architecture reduces the root mean square error (RMSE) to 0.48 t/ha with a coefficient of determination (R2) of 0.90. Statistical analysis proved the significant superiority of the developed method over industry-standard gap-filling algorithms (e.g., STARFM). A scalable decision-support system for optimizing harvest logistics and financial planning under any atmospheric conditions was presented. Future research directions involving neural network knowledge distillation for IoT devices and UAV data integration were outlined.
Why it matches plant phenotyping methods雲下で欠測するNDVI・NDREなど植物状態指標を再構成する深層学習手法の開発と、SSIM・RMSE・R2および既存手法との比較検証が中心であり、単なる収量予測ではなく植物表現型推定手法に該当する。
abstractDynamic synthesis of missing NDVI and NDRE vegetation indices was conducted under prolonged continuous cloud cover (up to 21 days), achieving a high structural similarity index (SSIM = 0.89).
Strawberry greenhouse cultivation is increasingly supported by sensing technologies, artificial intelligence (AI), and decision-support infrastructure, but their horticultural value depends on whether heterogeneous measurements can be translated into biologically meaningful crop states and practical management decisions. This review synthesizes strawberry phenotyping, multimodal sensing, AI-based crop-state interpretation, and supervised agentic coordination as a phenotyping-to-action framework for greenhouse strawberry cultivation. The reviewed studies show substantial progress in measuring and interpreting vegetative, reproductive, fruit-quality, stress-related, and environmental crop states through imaging, spectral, environmental, root-zone, and modeling approaches. However, much of the literature still emphasizes measurement accuracy, model performance, or infrastructure capability, whereas fewer studies validate whether AI-derived outputs improve crop response, management decisions, workflow, resource use, or production outcomes. The review therefore distinguishes sensing technologies for data acquisition and measurement from AI-based methods for interpretation and prediction, and examines how crop-state information can be connected to practical greenhouse decision making. It also compares established decision technologies, including expert systems, model predictive control, digital twins, and closed-loop coordination, with supervised agentic coordination as bounded decision-support concepts rather than as evidence of unrestricted autonomous control. Future work should emphasize phenotype-to-action validation, domain-aware benchmarking, and supervised deployment studies that connect model outputs with decision rules, crop outcomes, operational constraints, and grower oversight. By grounding sensing technologies and AI-based interpretation methods in crop-response validation, strawberry greenhouse systems can progress toward supervised, crop-state-driven decision support.
Why it matches plant phenotyping methods温室イチゴのフェノタイピング、マルチモーダルセンシング、AIによる作物状態解釈を中心に整理する方法論レビューであり、植物状態の取得・推定手法が主題。
abstractThis review synthesizes strawberry phenotyping, multimodal sensing, AI-based crop-state interpretation, and supervised agentic coordination as a phenotyping-to-action framework for greenhouse strawberry cultivation.
Modern agriculture operates at an unprecedented crossroads, it must simultaneously accelerate crop yields to feed an expanding global population and adapt to the severe, fluctuating pressures of climate change, structural soil degradation, abiotic water deficits, and evolving biological threats. Historically, selecting resilient crop varieties and implementing field-scale management strategies relied extensively on destructive, labor-intensive, and fundamentally subjective visual metrics. This manual processing approach has long been recognized as the primary operational bottleneck in agricultural advancement.To bridge the gap between rapidly expanding genomic data and actual field performance, the systematic, non-destructive quantification of structural and functional plant traits, plant phenotyping, has emerged as a transformative frontier. By integrating high-throughput engineering, multi-scale remote sensing, deep learning, and advanced molecular biology, modern phenotyping transitions crop science away from qualitative estimation toward highly reproducible, multidimensional data frameworks. This Research Topic presents new advances in advanced 3D reconstruction and deep semantic segmentation at the seedling stage; amodal fruit segmentation, morphological extraction, and early water-stress diagnostics; high-throughput in-field seedling counting and dynamic density modeling; multimodal foundation models, network pruning, and intelligent phytoprotection; aerial and spaceborne remote sensing for canopy analysis and weed monitoring; plant physiology, functional spectroscopy, and functional genomics under abiotic stress; and automated diagnostics for real-time orchard scouting and vineyard management.Automating the characterization of complex spatial layouts under controlled or greenhouse environments is essential for early variety selection and early-stage structural evaluation. Several contributions within this volume provide key breakthroughs in navigating overlapping tissues, severe occlusions, and low-contrast edge regions. showcases how substituting standard convolutions with deformable convolutions enables deep neural networks to accurately isolate the main stem of mature, high-density crops like soybeans. This architecture overcomes the traditional challenges of color mimicry and severe occlusion by pods and leaves, achieving an outstanding mIoU of 90.58% and providing reliable indices for lodging resistance and structural yield modeling (R 2 = 0.9746).Accurately extracting fruit morphology under commercial greenhouse conditions remains heavily constrained by overlapping crop structures, foliage cover, and variable shadows. Simple semantic masks typically fail when a target fruit is partially blocked, leading to a loss of key volumetric data.To resolve the challenge of hidden boundaries, Li, Yin, et al. (2025) developed CGA-ASNet, a specialized RGB-D amodal segmentation network driven by a Contextual and Global Attention (CGA) module designed to restore occluded tomato regions. Trained on a high-fidelity synthetic greenhouse dataset (Tomato-sim) generated via NVIDIA Isaac Sim's Replicator Composer and optimized with a mean coordinate fusion algorithm for real-world validation, this architecture expands the network's receptive field to predict the complete, hidden circular forms of occluded tomatoes, achieving an F@0.75 score of 94.2 and an amodal mIoU of 82.4%. This proves that simulation-to-real (Sim2Real) domain pathways can successfully decode full physical volumes under dense commercial canopies.Complementing this structural restoration, Yang, Li, et al. (2025) designed an integrated diagnostic framework to identify early water stress dynamics in greenhouse tomatoes. Built upon an optimized YOLOv11n core, their system integrates adaptive kernel convolutions (AKConv) into the network backbone's C3k2 modules and implements a recalibration feature pyramid detection head based on the specialized P2 small-target layer. This combination achieved a 5.4% increase in mAP50-95 for identifying fine phenotypic parts. By applying automated geometric analysis to the extracted bounding boxes, the system extracts plant heights and petiole count with low relative errors, feeding these phenotypic parameters into a Random Forest classification routine that flags water-stressed plants with 98% accuracy to guide targeted, automated drip irrigation.Accurate plant stands during early vegetative stages represent the foundational metric required to establish true field emergence rates, validate seed vigor across diverse breeding blocks, and perform early yield predictions.To solve the challenges of small targets, extreme spatial density, and adjacent leaf overlap, Zang et al. (2025) designed DM_IOC_fpn, a wheat seedling counting framework that balances local and global contextual features. By structuring a point-annotated dataset and embedding a densityenhanced encoder module, their network balances micro-scale spatial limits with macro-scale canopy structures. Optimized through a combined loss function tracking counting, classification, and regression parameters, this architecture achieved low error scores (RMSE = 2.91; MAE = 2.23), outperforming standard object-detection benchmarks in complex field environments.At the same time, scaling up to real-time aerial monitoring required major reductions in model complexity to support resource-constrained edge computers on autonomous aerial platforms. Feng, Nie, and Li (2025) engineered an ultra-lightweight YOLOv8n variant tailored for real-time maize seedling counting from high-speed UAV RGB overflights. By reparametrizing RepConv with HGNetV2, they constructed a lean Rep_HGNetV2 backbone, integrated a Bidirectional Feature Pyramid Network (BiFPN) for multi-scale feature alignment, and implemented a Task Dynamically Aligned Detection Head (TDADH). This architecture compressed total model parameters by 47% and reduced weight sizes to 3.5 MB while maintaining a 96.5% detection accuracy and an ultra-fast processing speed of 146.3 FPS, paving the way for low-cost, real-time field scouting.Automated phytoprotection requires machine-vision architectures capable of generalizing across highly diverse species, complex field conditions, and varying computational boundaries. A significant subset of the published papers addresses these challenges through foundation model adaptation, multi-modal alignment, and efficient network compression.A major paradigm shift presented in this collection involves moving away from task-specific training and toward foundation model adaptation. Chen, Ruan, et al. (2026) introduce a novel architecture integrating the DinoV3 foundation model with a Unet framework to achieve robust leaf lesion segmentation across diverse species (such as coffee and black gram). By incorporating a Spatial Prior Module (SPM), their approach surpassed standard benchmark networks by over 10.5% in IoU while reducing inference times by approximately 93.6%, demonstrating that highparameter foundation models can be highly optimized for resource-constrained edge devices in real-time scouting.To solve the perennial problem of limited training data for rare or emerging crop diseases, Cooper et al. ( 2026) developed an ingenious synthetic data generation pipeline. Combining 3D procedural leaf modeling in Blender with diffusion-based disease synthesis (Stable Diffusion fine-tuned with LoRA and ControlNet), they synthesized highly accurate plant disease images with perfect groundtruth annotation masks. When deployed in low-resource data settings, combining these synthetic pipelines with restricted real-world datasets consistently drives significant improvements in downstream segmentation tasks. To tackle specific, complex pathologies, Xu, Chang, et al. (2025) developed the TSSC deep learning model, which embeds three-neighbor channel attention paired with a complementary squeeze-and-excitation mechanism. This specific architecture minimizes structural degradation risks while pushing classification accuracy to 99.61% for highly complex pea leaf pathologies. Similarly, Feng, Liu, et al. (2025) tackled overlapping leaf occlusions and small lesion footprints in citrus groves with YOLO-Citrus, an optimized framework integrating C3K2-STA, ADown modules, and a Wise-Inner-MPDIoU loss function to strike a balance between edge computational constraints and field deployment.UAVs and high-resolution satellite imagery have expanded the operational scale of phenotyping from individual pots to vast breeding blocks and commercial fields, allowing researchers to capture macro-dynamic parameters over time.In complex canopy systems that defy standard top-down aerial sensing, such as single-staked white Guinea yams, Iseki et al. (2026) demonstrated the distinct advantage of utilizing multi-angle (combined nadir and oblique) UAV imaging configurations. When coupled with support vector regression, this method captures complementary canopy-structure information to model shoot biomass trajectories (R 2 = 0.79) across multiple years and management zones. These nondestructive, time-series datasets enabled the fitting of genotype-specific Richard's growth curves using Bayesian inference, isolating valuable genetic variations in early growth allocation.To capture full-season vertical physiological changes over large scales, Li, Yue, and Luo (2025) developed a hybrid CNN-LSTM-Attention (CLA) model designed to estimate the full-period Leaf Area Index (LAI) in rice using multi-temporal UAV multispectral imagery. By using the CNN layer to extract instantaneous spatial features, the LSTM block to process seasonal time-series intervals, and a self-attention mechanism to weight critical growth transitions, their platform achieved a high coefficient of determination (R 2 = 0.92) and kept relative root mean square errors (RRMSE) below 9%. This network minimized soil background noise during early vegetative stages (LAI values 1-
Why it matches plant phenotyping methods植物フェノタイピングの技術動向を扱うEditorialであり、画像解析、UAVセンシング、深層学習、形質抽出などの方法が中心的に整理されている。
Olive is a major agricultural crop extensively cultivated throughout the Mediterranean region. However, olive trees are vulnerable to several diseases that can negatively affect productivity and yield. One of the most widespread foliar diseases is olive leaf peacock spot, caused by the fungus Cycloconium oleaginum. Early detection of this disease is essential for preventing leaf drop, limiting disease spread, maintaining tree health, and reducing treatment costs before the infection reaches an advanced stage. In this study, a multimodal hybrid deep learning framework is developed to detect peacock spot disease in olive leaves and assess disease severity based on visual and numerical features. The proposed framework integrates olive leaf images with soil conditions, environmental conditions, and vegetation and stress indices to provide a more comprehensive disease analysis than image-only approaches. A ResNet50-based convolutional neural network is used to extract visual features from leaf images, while a multilayer perceptron processes the numerical sensor-based and index-based data. These features are then fused within a unified learning framework to classify disease stages and estimate leaf damage severity, including lesion coverage and yellowing percentage. The performance of the proposed model was evaluated using standard performance metrics suitable for both classification and regression tasks. For classification, the model was evaluated on 494 testing samples and achieved an overall accuracy of 97.77 %, with a macro F1-score of 0.9809 and a weighted F1-score of 0.9776. In addition, the model achieved low regression errors, with mean absolute errors of 1.16 % for lesion coverage and 1.42 % for yellowing estimation. These results demonstrate the effectiveness of the proposed multimodal framework for accurate peacock spot detection and severity assessment, supporting its potential use in smart agricultural monitoring and disease management.
Why it matches plant phenotyping methods画像とセンサーデータを統合し、オリーブ葉の病斑被覆率・黄化率という植物の病害状態を推定する深層学習手法が研究の中心であるため。
abstracta multimodal hybrid deep learning framework is developed to detect peacock spot disease in olive leaves and assess disease severity based on visual and numerical features
Crop diseases play a significant role in food production globally; therefore, there is an urgent need to develop quick and accurate diagnostic techniques that are more effective than manual inspection methods. The proposed hybrid multimodal learning framework in this research provides a solution that integrates adaptive therapy suggestion, market price prediction, and image-based disease detection. This study also proposes a framework for pesticide recommendation and the treatment of plants. This study experiment on tomato and cotton crop leaf data for disease detection. Experimental results on a tomato crop disease detection dataset show that the proposed model shows high performance. EfficientNetB0 provides more stability and generalization capabilities in different scenarios compared to other models, such as YOLOv8, ResNet50, and a custom CNN model. The use of a knowledge-based decision support system provides sustainable pesticide recommendations based on environmental and symptom-specific parameters. Forecasting of pesticide prices through LSTM methods yields forecasts within 3.2% and 4.1% MAE, enabling improved decision-making by providing instant points of reference for potential price movements. Research uses SHAP and LIME to provide explainability to users, thus improving user buy-in through transparency. Overall, this modular system provides a data-driven decision-making model to improve the efficiency of managing crops.
Why it matches plant phenotyping methods植物葉画像から病害状態を推定する画像ベース手法を、複数モデルで比較評価しており、植物病害フェノタイピングがシステムの主要構成要素です。価格予測や農薬推薦も含みますが、病害検出の技術評価が明示されています。
abstractThe proposed hybrid multimodal learning framework in this research provides a solution that integrates adaptive therapy suggestion, market price prediction, and image-based disease detection.
Reproduction assets foundThe paper's disease-detection experiments use publicly available cotton and tomato leaf image datasets (Kaggle, IEEE DataPort, Roboflow), all cited with explicit public URLs in the references. No author analysis code or trained model checkpoints are stated as publicly available; the supplementary material is referencedDataset · publiccholar
View reference in article
19
Muppala C. Guruviah V. ( 2020 ). Machine vision detection of pests, diseases, and weeds: a review . J. Phytol. 12 , 9 – 19 . doi: 10.25081/jp.2020.v12.6145
CrossRef
Google Scholar
View reference in article
20
National College of Ireland ( 2025 ). “Cotton Disease Dataset.” Available online at: https://www.kaggle.com/datasets/janmejaybhoi/cotton-disease-dataset (Accessed May 19, 2025).
Google Scholar
View reference in article
21
Naveed Gul and Kaggle ( 2026 ). Tomato Leaf Disease . Kaggle. Available online at: https://www.kaggle.com/datasets/naveedgull/tomato-leaf-disease (Accessed March 29, 2026).
Google Scholar
View reference in article
22
Ngugi H. N. EzugOpen asset ↗Kagglelines:554-633Dataset · publicreference in article
20
National College of Ireland ( 2025 ). “Cotton Disease Dataset.” Available online at: https://www.kaggle.com/datasets/janmejaybhoi/cotton-disease-dataset (Accessed May 19, 2025).
Google Scholar
View reference in article
21
Naveed Gul and Kaggle ( 2026 ). Tomato Leaf Disease . Kaggle. Available online at: https://www.kaggle.com/datasets/naveedgull/tomato-leaf-disease (Accessed March 29, 2026).
Google Scholar
View reference in article
22
Ngugi H. N. Ezugwu A. E. Akinyelu A. A. Abualigah L. ( 2024 ). Revolutionizing crop disease detection with computational deep learning: a comprehensive review . Environ. Monit. Assess. 196 : 302 . doi: 10.1007/s10661-024-12454-z
Pubmed AOpen asset ↗Kagglelines:554-633Dataset · publicComputer Vision and Pattern Recognition (CVPR) ( Las Vegas, NV : IEEE ), 779 – 788 . doi: 10.1109/CVPR.2016.91
CrossRef
Google Scholar
View reference in article
29
Roboflow ( 2026a ). A Comprehensive Dataset of Cotton Plant Diseases for National Disease Identification and Treatment Guidance | IEEE DataPort. Available online at: https://ieee-dataport.org/documents/comprehensive-dataset-cotton-plant-diseases-national-disease-identification-and-treatment (Accessed March 29, 2026).
Google Scholar
View reference in article
30
Roboflow ( 2026b ). Cotton Plant Disease Prediction Object Detection Model by National College of Ireland . Available online at: https://universe.roboflow.com/national-colleOpen asset ↗IEEE DataPortlines:554-633Plant phenotyping relevance match · UnverifiedEurope PMC · checked 5 Sept 2026
In complex mountainous environments, the asynchronous development between external color turning and internal sugar accumulation (often termed "false maturity") in coffee cherries poses a severe challenge to post-harvest quality sorting and the consistency of final coffee products. To overcome the limitations of single-phenotype detection in raw material screening, this study proposed a multimodal quality discrimination framework integrating fruit hyperspectral imaging, micro-topography, and plant physiological characteristics. Taking typical mountain-grown fresh coffee cherries as the research object, and after comparing various spectral preprocessing and feature dimensionality reduction algorithms, the multimodal fusion efficacy of nine machine learning classifiers was systematically evaluated. The results demonstrated that: (1) Full-spectrum difference analysis quantitatively confirmed the limitations of visual harvesting; spectral reflectance differences between high- and low-sugar fruits were highly concentrated in the red and red-edge regions, with the maximum difference precisely located at 676 nm. (2) Compared to the single-spectrum model (mean accuracy of 75.93%), the fully fused Multilayer Perceptron (MLP) network effectively mitigated background noise induced by heterogeneous environments, improving the mean classification accuracy to 77.22% with a mean Area Under the Curve (AUC) of 0.827. (3) Correlation analysis clarified the quantitative association between topography and quality; micro-topographic slope (r = 0.346) was identified as the key environmental driver of spatial differentiation in fruit sugar content, while plant chlorophyll A content (r = 0.183) exhibited a corresponding physiological response trend. This study not only explains the root cause of visual assessment failure from a physical optics perspective but also reveals the spatial variation laws of quality driven by micro-topography, providing preliminary data support for the intelligent sorting of raw materials and ensuring post-harvest quality consistency of mountainous crops.
Why it matches plant phenotyping methodsコーヒー果実の糖含量という植物器官形質を、ハイパースペクトル画像・微地形・生理情報の融合で非破壊推定する方法が研究の中心であり、前処理、特徴削減、複数分類器の性能比較も行っている。
abstracta multimodal quality discrimination framework integrating fruit hyperspectral imaging, micro-topography, and plant physiological characteristics
Despite the widespread use of image-only convolutional models for plant disease diagnosis to provide global food security, image variability at the field level, visually similar symptoms, and stress due to soil nutrients or moisture conditions, can cause loss of accuracy. This paper compares classical machine learning classification models to custom convolutional neural networks and pretrained transfer-learning models for multi-crop and multi-disease classification and also introduces a late-fusion multimodal decision support model that could integrate data about images, soil, and weather. Experiments were conducted on the PlantVillage dataset which had 54,305 RGB images belonging to 38 different crop-condition classes (43,456 training, 10,849 validation, and 10,849 test images) across 14 different crops. The accuracy, macro-precision, macro-recall and macro-F1 score were computed for the five classical classifiers (KNN, Random Forest, Extra Trees, SGD-linear SVM, and SVC-RBF using PCA-reduced features), three custom CNN variants, and four pretrained models (ResNet50, MobileNetV2, GoogleNet, and EfficientNetB7). Of the classical models, SVC with RBF kernel yielded the highest accuracy (83.58%) and MacroF1 (82.96%). In the case of the custom CNN models, the deeper they became and the more dropout the higher accuracy they gained, with CNN-V3 attaining 95.71% accuracy and 95.66% macro-F1. Transfer learning yielded the best results with MobileNetV2 (99.30% accuracy and macro-F1) outperforming ResNet50 (98.60% accuracy and macro-F1) and GoogleNet (97.83% accuracy and 97.46% macro-F1). The findings serve as the basis for the introduction of a multimodal system for soil forecasting integrating a MobileNetV2 image encoder, a residual MLP for soil features, and two dual LSTMs for 48-hour and 168-hour time-series of the weather. The interpretability, modularity and the sufficient resistance towards loss of sensor information for probability-level late fusion with 0.80, 0.10, 0.05 and 0.05 respectively make the framework appropriate for precision-agriculture advisory systems, until it is tested in the field.
Why it matches plant phenotyping methods植物画像から病害状態を推定する分類手法を複数比較し、マルチモーダル意思決定フレームワークを構築・評価しており、病害表現型の取得・推定が中心である。
abstractThis paper compares classical machine learning classification models to custom convolutional neural networks and pretrained transfer-learning models for multi-crop and multi-disease classification
A hybrid AI framework combining a spatial–spectral–temporal Transformer and unsupervised clustering was applied to five microgreen species—Pak Choi Cabbage, Tatsoi Mustard, Red Mizuna, Chinese Cabbage, and Arugula—grown for 3–4 weeks under water, nutrient, and combined stresses. Across five datasets collected within six months, the system achieved a Macro-F1 of 0.91 and a pre-symptomatic F1 of 0.88, enabling early detection before visible symptoms. Applications include NASA’s APH, Mars and Moon habitats, and terrestrial precision agriculture.
Why it matches plant phenotyping methods植物の水・養分ストレス状態をマルチモーダル画像から早期推定するAI手法が研究の中心であり、性能評価も示されている。
abstractA hybrid AI framework combining a spatial–spectral–temporal Transformer and unsupervised clustering was applied
The detection technology of crop diseases and pests is transitioning from single sensor monitoring to intelligent perception and multimodal fusion. This paper follows the PRISMA 2020 standard and systematically reviews the relevant core literature. This paper systematically summarizes the development history of spectral sensing technology and analyzes the physical mechanisms of hyperspectral and multispectral imaging in early identification of crop diseases. The focus is on the architectural evolution of deep learning models, including lightweight convolutional neural networks (CNNs), vision transformers (ViTs) with long-range dependency modeling capabilities, and the efficient computing state space model Mamba. In addition, the research progress of spatial spectral joint learning, heterogeneous data fusion, and vision-language models (VLMs) in improving system robustness and interpretability are introduced. By synthesizing the integrated applications of UAV remote sensing, Internet of Things (IoT) edge computing and intelligent robots in staple and cash crops, this paper summarizes the implementation of the integrated system of perception, decision-making and execution. To address the issues of insufficient cross-domain generalization ability and uneven allocation of computing resources in existing models, this paper provides perspectives on the future development of agricultural artificial intelligence (AI) towards foundation model-driven, edge-intelligent collaboration, and green sustainable direction, which can provide theoretical reference for engineering applications in the field of intelligent plant protection.
Why it matches plant phenotyping methods作物病害の画像・スペクトル観測による植物の病徴・病害状態推定を中心に、検出技術と計算手法を体系的にレビューしており、植物フェノタイピング手法のレビューに該当する。
abstractThis paper systematically summarizes the development history of spectral sensing technology and analyzes the physical mechanisms of hyperspectral and multispectral imaging in early identification of crop diseases.
Accurate tree counting from remote sensing data is essential for forest inventory, biomass estimation, carbon accounting, and ecological monitoring. However, existing approaches predominantly rely on airborne RGB imagery and often struggle in complex forest scenes where neighboring crowns exhibit highly similar textures and colors and where overlapping crown boundaries become ambiguous. To address this limitation, the LiDAR-derived Canopy Height Model (CHM) is introduced as a complementary modality that provides explicit cues on canopy height variation and vertical structure to support RGB-based analysis. Building on this, we propose BCAR-Net, a broker-guided RGB and depth (RGB-D) multimodal framework that couples bidirectional cross-modal interaction, adaptive tri-branch fusion, and auxiliary reconstruction within a two-stage optimization scheme. Specifically, a bidirectional cross-attention U-Net generates an intermediate broker RGB-D representation from paired RGB images and depth maps through symmetric bidirectional cross-attention between the two modalities and direction-aware gating. The original RGB image, depth map, and broker representation are then jointly encoded by three weight-sharing branches and adaptively aggregated by a spatial fusion gate for density-map regression. To regularize the fused latent feature, a multi-scale cross-attention reconstruction decoder provides auxiliary RGB and depth reconstruction supervision by querying multi-scale BCA-UNet encoder features through 2D cross-attention, and a reconstruction-oriented first stage replaces externally generated fused-image supervision, yielding a task-consistent optimization scheme. Experiments on the NEONTreeEvaluation benchmark show that BCAR-Net consistently outperforms single-modality settings and direct RGB-D concatenation multimodal baseline. Additional experiments on a public UAV RGB-LiDAR dataset provide a small-scale supplementary evaluation under a different acquisition setting, where BCAR-Net achieves modest but consistent improvements over RGB-only and depth-only baselines. These results demonstrate that the proposed framework offers an effective but computationally cautious solution for tree counting in complex forest environments.
Why it matches plant phenotyping methodsRGB画像とLiDAR由来データから樹木数を推定する深層学習手法を開発し、複数ベンチマークで比較評価しており、植物個体の計測手法が研究の中心である。
abstractwe propose BCAR-Net, a broker-guided RGB and depth (RGB-D) multimodal framework
The high-precision instance segmentation of tree saplings is a fundamental prerequisite for the high-throughput phenotypic analysis of individual seedlings in intelligent tree breeding and precision silviculture. However, sapling segmentation remains challenging because of blurred boundaries, object adhesion, missed detections, and inaccurate mask delineation in field environments. To improve sapling segmentation performance and address these challenges, this study proposes a multimodal Mask R-CNN framework in which RGB imagery was paired with one multispectral-derived vegetation index at a time to construct separate RGB-VI input combinations, taking ginkgo saplings as a representative case. A dataset of 400 saplings was constructed using a high-throughput field phenotyping platform. The backbone network was extended with an independent vegetation index branch, and three fusion strategies (early, multi-step, and late fusion) were designed within a feature pyramid network to enable multi-scale multimodal feature integration. The results showed that all multimodal models outperformed unimodal baselines in terms of segmentation accuracy and recall. Among them, the multi-step fusion strategy achieved the best performance, while the RGB-EVI multi-step fusion model achieved the highest strict-matching precision (AP@75 = 87.7%) and recall (71.3%), with superior performance in dense sapling delineation and background suppression. These findings indicate that multimodal feature fusion can effectively improve sapling instance segmentation and provide methodological support for high-throughput plant phenotyping.
Why it matches plant phenotyping methodsマルチモーダル画像による樹木苗個体のインスタンスセグメンテーション手法を開発・比較し、高スループット表現型解析を支援することが中心である。
abstractThe high-precision instance segmentation of tree saplings is a fundamental prerequisite for the high-throughput phenotypic analysis of individual seedlings in intelligent tree breeding and precision silviculture.
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
Abstract Deep learning, as a pivotal branch of machine learning, has demonstrated remarkable potential in advancing crop science by effectively integrating genomics and phenomics. This review systematically outlines the application of diverse deep learning architectures—such as convolutional neural networks, recurrent neural networks, and transformers—across key crop genomic tasks, including gene expression prediction, alternative splicing analysis, cis ‐regulatory element identification, epigenomic profiling, and genome‐based trait prediction. In phenomics, these models facilitate high‐throughput extraction of crop phenotypic traits from multispectral, unmanned aerial vehicle, and ground‐based imagery, supporting yield forecasting, disease diagnosis, and stress response monitoring. We critically evaluate the performance and limitations of each model type across tasks, considering trade‐offs between complexity, accuracy, and interpretability, to offer practical guidance for crop researchers. Additionally, the review addresses major challenges in deploying deep learning—such as data scarcity, model transparency, and computational demands—and proposes future pathways to enhance model generalizability, multimodal data integration, and applications in intelligent breeding and sustainable agriculture.
Why it matches plant phenotyping methods作物フェノミクスにおける深層学習による画像からの形質抽出を中心的にレビューしており、フェノタイピング手法の方法論的整理に該当する。
abstractIn phenomics, these models facilitate high‐throughput extraction of crop phenotypic traits from multispectral, unmanned aerial vehicle, and ground‐based imagery
Accurate aboveground biomass estimation with quantified uncertainty is essential for precision agriculture, enabling risk-aware decision-making and strategic model improvement. Existing approaches predominantly provide point estimates without uncertainty quantification, limiting their operational utility for trustworthy Artificial Intelligence (AI) deployment. This study presents a Multi-modal Attention-based Uncertainty Quantification Network (MA-UQNet), which achieves superior prediction accuracy (R 2 = 0.856) with well-calibrated uncertainty (97.18% coverage) for wheat aboveground biomass estimation through integrated multi-modal attention, growth stage-specific processing, and epistemic–aleatoric uncertainty decomposition. The framework integrates hyperspectral remote sensing with environmental variables via joint attention mechanisms that adapt to phenological variations. Model development employed a decade-spanning dataset (2012–2022, 1272 samples) collected under factorial combinations of nitrogen rates (0–270 kg/ha), irrigation levels (0–384 mm), and wheat cultivars across four growth stages. Temporal extrapolation validation using chronological partitioning (2012–2019 for training and 2020–2021 for testing) demonstrated robust generalization, substantially outperforming Random Forest (R 2 = 0.751, coverage = 76.61%) and nine representative baselines, including Bayesian Neural Networks (R 2 = 0.805, coverage = 38.31%). Uncertainty decomposition revealed epistemic uncertainty to be moderately dominant (53%) relative to aleatoric uncertainty (47%), indicating that strategic data collection offers greater potential for uncertainty reduction than improving measurement precision alone. These findings provide validated tools for uncertainty-aware biomass estimation in precision agriculture.
Why it matches plant phenotyping methods小麦の地上部バイオマスという植物形質を、ハイパースペクトルリモートセンシングと不確実性推定ネットワークで抽出する手法を開発し、時系列分割と既存手法との比較で検証しているため、植物フェノタイピング手法が中心である。
abstractThis study presents a Multi-modal Attention-based Uncertainty Quantification Network (MA-UQNet), which achieves superior prediction accuracy (R 2 = 0.856) with well-calibrated uncertainty (97.18% coverage) for wheat aboveground biomass estimation
This paper presents a novel multimodal deep learning framework for maple plant disease detection by integrating visual and semantic information. Traditional plant disease detection systems rely primarily on visual features extracted from leaf images, which often leads to misclassification in cases of visually similar disease symptoms. To address this limitation, the proposed approach combines EfficientNet-based convolutional neural networks for visual feature extraction with transformerbased language models, including BERT and FLAN-T5, for semantic feature encoding. A Multilayer Perceptron (MLP)-based fusion mechanism is employed to integrate visual and textual features, enabling effective cross-modal learning. The proposed model is evaluated on a balanced dataset of 2,000 maple leaf images and associated disease descriptions. Experimental results demonstrate that the multimodal framework achieves an accuracy of 94.8%, outperforming vision-only and text-only models by a significant margin. Ablation studies and comparative analysis confirm the effectiveness of multimodal fusion and transformerbased semantic encoding. The proposed framework provides a robust and scalable solution for intelligent plant disease detection and has potential applications in smart agriculture systems.
Why it matches plant phenotyping methods葉画像から植物病害状態を推定するマルチモーダル画像解析手法を開発・比較評価しており、植物フェノタイピング手法が中心である。
abstractThis paper presents a novel multimodal deep learning framework for maple plant disease detection by integrating visual and semantic information.
Despite numerous efforts to incorporate emerging innovations into the agricultural domain for increasing crop yield and actively managing the state of the fields, it remains difficult for the industry to implement cutting-edge technologies in practice. This paper proposes AgroVision – a web-based intelligent multimodal system for comprehensive analysis of crops around the world. Designed using a three-layer scalable architecture, the system includes four modules – CNN-based growth stages and plant diseases recognition, AI Chatbot with LLM capabilities and RAG support, as well as the video analysis tool for detecting plant density and weeds. The key technology behind the core image analysis functionality of AgroVision is represented by the efficient Vision Mamba (ViM) architecture, which allows for analysing multiple tasks simultaneously using only one image uploaded by the user. Based on the extensive dataset called "New Plant Diseases Dataset" containing over 87 thousand images divided into 38 classes, the ViM model demonstrates exceptional results achieving weighted average F1-Score of 97.1%. Considering that the inference latency of the model does not exceed 25-40 milliseconds, the system can be deployed at the edge, providing an easy-to-use solution for farmers.
Why it matches plant phenotyping methods植物の成長段階、病害、植物密度を画像から推定するウェブ型解析プラットフォームを提案しており、表現型取得・推定が研究の中心である。
abstractThis paper proposes AgroVision – a web-based intelligent multimodal system for comprehensive analysis of crops around the world.
The utilization of multi-source sensing data to achieve intelligent perception and refined management of farmland has become a vital research direction in modern agriculture. However, traditional inspection approaches based solely on visual information are highly susceptible to illumination variations, occlusion, and background interference, which makes stable pest detection and accurate crop growth assessment difficult to achieve. To address these problems, we propose a multimodal target perception network for intelligent farmland inspection. By integrating UAV imagery, ground environmental sensor data, and spatial location information, joint perception of farmland pests, diseases, and crop growth status is achieved. In the proposed framework, cross-modal alignment and collaborative encoding mechanisms, a multi-scale target perception structure, and a dynamic multimodal fusion strategy are introduced to collaboratively model information within a unified semantic space. Experimental results on a constructed multimodal farmland dataset demonstrate that the proposed method achieved 87.53% Precision and 89.16% mAP in the pest and disease detection task, and 88.04% Accuracy in the crop growth assessment task, significantly outperforming several mainstream visual detection models and multimodal fusion approaches. The results indicate that this intelligent perception framework can significantly improve the robustness of farmland inspection systems, providing an effective technical pathway for AI-driven precision agriculture decision-making. This technology breaks the barrier between production-side sensing data and e-commerce demand, providing a practical technical solution for agricultural production-marketing synergy, quality premium realization and digital rural revitalization.
Why it matches plant phenotyping methodsUAV画像と環境センサーを統合し、作物の生育状態および病害を推定するマルチモーダル手法を開発・評価しており、植物表現型の取得が中心的な貢献である。
abstractBy integrating UAV imagery, ground environmental sensor data, and spatial location information, joint perception of farmland pests, diseases, and crop growth status is achieved.
Effective crop health monitoring requires the integration of heterogeneous data sources that capture both environmental conditions and crop-level responses. Conventional single-modality approaches, relying either on in-ground sensor measurements or aerial imagery in isolation, fail to exploit the complementary strengths of each technique, resulting in limited diagnostic accuracy and delayed intervention. This study proposes an integrated approach that fuses data acquired by IoT-connected sensor networks with multispectral images captured by unmanned aerial vehicles (UAVs), enabling comprehensive crop health assessment across an entire field. Consider a system that deploys a network of sensors throughout a 2-hectare area to continuously watch the moisture and temperature, as well as other critical parameters, such as electrical conductivity, atmospheric humidity as well as the nutrient level in the soil. Simultaneously, a special UAV, a drone known as the DJI Phantom 4 Multispectral, take pictures of the field at five different spectral frequencies. However, the most interesting thing is the following: to do all this, the system relies on a special type of artificial intelligence known as a hybrid deep learning architecture. It resembles a twostep procedure, where the one section analyses the trends in the sensor data across time and is called one dimensional convolutional neural network, whereas the other section uses a pre-trained version of ResNet-50 and is responsible of making significant features of the images captured by the drone. Then, it is all unified with an eight-head multi-head attention mechanism, taking all the various kinds of data and forming a bigger picture. This will enable the system to memorize and establish relationships among the various kinds of data, which forms a potent source of knowledge and management of the field. The synchronized dataset comprised 2,520 data points collected over 120 days (April–August 2024), with 103 days of active data capture. The proposed hybrid deep learning architecture achieved a crop health classification accuracy of 94.5%, compared to 80% for conventional single-modality methods — a statistically significant improvement of 14.5 percentage points. The system classifies crop status into five categories: healthy, water stress, nutrient deficiency, disease, and pest infestation. The results demonstrate that cross-modal IoT–image fusion delivers earlier, more reliable diagnosis of crop stress conditions, enabling data-driven farm management decisions that support precision agriculture at scale.
Why it matches plant phenotyping methodsUAVマルチスペクトル画像とIoTセンサーデータを融合し、作物の健康状態・ストレス状態を分類する深層学習手法が研究の中心であり、植物状態の取得・推定方法として適格。
abstractThis study proposes an integrated approach that fuses data acquired by IoT-connected sensor networks with multispectral images captured by unmanned aerial vehicles (UAVs), enabling comprehensive crop health assessment across an entire field.
Agricultural crop diseases can greatly reduce production quality and overall farm output, making early identification important for sustainable farming. This study introduces a smart agricultural rover that applies a multimodal deep learning approach for real-time crop disease monitoring in field environments. The proposed system gathers RGB images, thermal information, and environmental measurements such as temperature, humidity, and soil moisture through integrated sensors connected to a Raspberry Pi 4. For on-device analysis, a lightweight TensorFlow Lite (TFLite) model is utilized to classify crop diseases efficiently at the edge. To improve detection performance under different illumination conditions, the system evaluates both original and CLAHE-enhanced images using a dualinference mechanism supported by entropy and confidence-based decision metrics. The rover is implemented on a mobile robotic platform equipped with motor control and battery support to enable autonomous movement in agricultural fields. By combining sensor fusion, edge intelligence, and robotic mobility, the developed system supports accurate identification of diseases such as Powdery Mildew and Rust, helping farmers take preventive action and improve crop management practices
Why it matches plant phenotyping methodsRGB・熱画像とセンサ融合、エッジ推論、画像強調による作物病害検出システムを開発しており、植物の病害状態を推定する方法が中心である。
abstractThis study introduces a smart agricultural rover that applies a multimodal deep learning approach for real-time crop disease monitoring in field environments.
To address the poor regression performance caused by strong spatiotemporal heterogeneity, inconsistent information scales and complex feature relationships in field-based multimodal data, this study proposes a hierarchical fusion framework embedded within an encoder – decoder network. The framework integrates multi-scale, interpretable, and cross-modal representations through coordinated modules that bridge feature discrepancy understanding and information flow regulation. A multi-scale parallel pathway structure is designed to enhance joint perception of local and global information by leveraging feature mappings with different receptive fields. An interpretable feature importance allocation strategy is further introduced to improve the backbone network’s ability to provide dynamic guidance on feature contributions. This enables the model to perform adaptive weighting and feature selection during multimodal fusion. In addition, a cross-modal dense interaction and gated fusion mechanism is constructed to regulate information flow and capture fine-grained associations across modalities. The improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties. Results show that, in the first year, the highest prediction accuracy reaches an R2 of 0.8112 with an rRMSE of 16.85%. Validation using data from the same planting region in the second and third years yields a highest R2 of 0.8107 and 0.7986, with corresponding rRMSE values of 17.92% and 17.27%, respectively. Compared with other deep learning models within the same year, the proposed approach improves R2 by up to 30.52%, 33.96% and 32.79% across the three years, while reducing rRMSE by up to 41.59%, 45.01% and 46.43%. The results demonstrate that the coordinated interaction among modules establishes an integrated optimization pathway that spans from feature discrepancy understanding to information flow regulation, while maintaining interpretability in the decision process. Under complex field conditions with multiple sources of uncertainty, the proposed framework achieves stable module contributions ranging from 5% to 10% based on cross-validation and t-test analyses. This effectively alleviates the difficulty of efficient multimodal feature fusion for robust yield prediction under stress conditions. The study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings.Develops a weight-guided method to enhance interpretability of multimodal features.Designs a cross-modal dense interaction and gated fusion mechanism.Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings. Develops a weight-guided method to enhance interpretability of multimodal features. Designs a cross-modal dense interaction and gated fusion mechanism. Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features.
Why it matches plant phenotyping methods小麦の収量という植物形質を対象に、マルチモーダルデータ融合と収量回帰のための新規エンコーダ・デコーダ手法を開発し、複数年・品種で検証している。フェノタイピング手法が中心である。
abstractthis study proposes a hierarchical fusion framework embedded within an encoder – decoder network
To address the poor regression performance caused by strong spatiotemporal heterogeneity, inconsistent information scales and complex feature relationships in field-based multimodal data, this study proposes a hierarchical fusion framework embedded within an encoder – decoder network. The framework integrates multi-scale, interpretable, and cross-modal representations through coordinated modules that bridge feature discrepancy understanding and information flow regulation. A multi-scale parallel pathway structure is designed to enhance joint perception of local and global information by leveraging feature mappings with different receptive fields. An interpretable feature importance allocation strategy is further introduced to improve the backbone network’s ability to provide dynamic guidance on feature contributions. This enables the model to perform adaptive weighting and feature selection during multimodal fusion. In addition, a cross-modal dense interaction and gated fusion mechanism is constructed to regulate information flow and capture fine-grained associations across modalities. The improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties. Results show that, in the first year, the highest prediction accuracy reaches an R2 of 0.8112 with an rRMSE of 16.85%. Validation using data from the same planting region in the second and third years yields a highest R2 of 0.8107 and 0.7986, with corresponding rRMSE values of 17.92% and 17.27%, respectively. Compared with other deep learning models within the same year, the proposed approach improves R2 by up to 30.52%, 33.96% and 32.79% across the three years, while reducing rRMSE by up to 41.59%, 45.01% and 46.43%. The results demonstrate that the coordinated interaction among modules establishes an integrated optimization pathway that spans from feature discrepancy understanding to information flow regulation, while maintaining interpretability in the decision process. Under complex field conditions with multiple sources of uncertainty, the proposed framework achieves stable module contributions ranging from 5% to 10% based on cross-validation and t-test analyses. This effectively alleviates the difficulty of efficient multimodal feature fusion for robust yield prediction under stress conditions. The study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings.Develops a weight-guided method to enhance interpretability of multimodal features.Designs a cross-modal dense interaction and gated fusion mechanism.Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features. Proposes a multi-scale parallel-path architecture for joint perception of multimodal feature mappings. Develops a weight-guided method to enhance interpretability of multimodal features. Designs a cross-modal dense interaction and gated fusion mechanism. Establishes an encoder–decoder-based framework for coordination and fusion of heterogeneous multimodal features.
Why it matches plant phenotyping methodsマルチモーダル作物センシングから小麦収量を推定する encoder–decoder 手法を開発し、複数年・品種データで検証しており、表現型推定手法が研究の中心である。
abstractthis study proposes a hierarchical fusion framework embedded within an encoder – decoder network.
To address the poor regression performance caused by strong spatiotemporal heterogeneity, inconsistent information scales and complex feature relationships in field-based multimodal data, this study proposes a hierarchical fusion framework embedded within an encoder – decoder network. The framework integrates multi-scale, interpretable, and cross-modal representations through coordinated modules that bridge feature discrepancy understanding and information flow regulation. A multi-scale parallel pathway structure is designed to enhance joint perception of local and global information by leveraging feature mappings with different receptive fields. An interpretable feature importance allocation strategy is further introduced to improve the backbone network’s ability to provide dynamic guidance on feature contributions. This enables the model to perform adaptive weighting and feature selection during multimodal fusion. In addition, a cross-modal dense interaction and gated fusion mechanism is constructed to regulate information flow and capture fine-grained associations across modalities. The improved feature-guided model is applied to yield regression using multimodal data collected over three consecutive years and multiple wheat varieties. Results show that, in the first year, the highest prediction accuracy reaches an R2 of 0.8112 with an rRMSE of 16.85%. Validation using data from the same planting region in the second and third years yields a highest R2 of 0.8107 and 0.7986, with corresponding rRMSE values of 17.92% and 17.27%, respectively. Compared with other deep learning models within the same year, the proposed approach improves R2 by up to 30.52%, 33.96% and 32.79% across the three years, while reducing rRMSE by up to 41.59%, 45.01% and 46.43%. The results demonstrate that the coordinated interaction among modules establishes an integrated optimization pathway that spans from feature discrepancy understanding to information flow regulation, while maintaining interpretability in the decision process. Under complex field conditions with multiple sources of uncertainty, the proposed framework achieves stable module contributions ranging from 5% to 10% based on cross-validation and t-test analyses. This effectively alleviates the difficulty of efficient multimodal feature fusion for robust yield prediction under stress conditions. The study provides a new methodological perspective for multimodal agricultural sensing and crop phenotyping.
Why it matches plant phenotyping methodsマルチモーダル農業センシングから小麦収量という植物形質を推定するエンコーダ・デコーダ手法を開発し、複数年・品種データで検証しているため、フェノタイピング手法が中心的である。
abstractthis study proposes a hierarchical fusion framework embedded within an encoder – decoder network.
The accurate quantification of glucoraphanin (GRA), a crucial health-promoting compound in broccoli, is vital for assessing its nutritional quality. However, traditional methods relying on destructive laboratory assays hinder rapid quality monitoring. To address this limitation, we developed a novel non-destructive, multimodal deep learning framework that integrates two phenotypic data modalities—image-based phenotypes from red-green-blue (RGB) leaf images and field-measured plant morphological traits—for accurate GRA estimation. Our proposed model, Parallel-Enhanced FasterNet (PE-FasterNet), incorporates two key innovations: a Gated Parallel Routing Attention (GPRA) mechanism for enhanced feature extraction, and a Phenotype-Guided Cross-Attention Feature Fusion (PG-CAFF) module for effective cross-modal fusion. Through rigorous evaluation, the model achieved a standard random-split test R 2 of 0.985 and a Leave-One-Group-Out (LOGO) cross-validation R 2 of 0.979, demonstrating highly accurate and generalized GRA predictions. This performance represents a substantial improvement over state-of-the-art convolutional neural network (CNN) and Vision Transformer models, affirming the architectural superiority of our approach. This study not only provides a robust tool for rapid, non-destructive prediction of GRA but also demonstrates a viable pathway toward data-driven crop quality management and precision breeding in broccoli.
Why it matches plant phenotyping methodsブロッコリー葉画像と形態形質からグルコラファニンを非破壊推定する深層学習法を開発・検証しており、表現型取得・抽出ワークフローが研究の中心である。
abstractwe developed a novel non-destructive, multimodal deep learning framework that integrates two phenotypic data modalities—image-based phenotypes from red-green-blue (RGB) leaf images and field-measured plant morphological traits—for accurate GRA estimation.
Plant phenotyping relevance match · UnverifiedCrossref · OpenAlex · Europe PMC · checked 5 Sept 2026
Abstract Climate change increasingly threatens global agriculture by intensifying abiotic stresses and destabilizing crop productivity, necessitating a deeper understanding of root-mediated traits governing resource acquisition and stress resilience. Here, we synthesize recent advances in root-centred plant phenomics, emphasizing how high-throughput phenotyping enables high-resolution, scalable characterization of complex root traits and robust comparative analysis across diverse genotypes and environments. Innovations in multimodal imaging, notably X-ray computed tomography, MRI, and machine learning-integrated rhizotrons, facilitate detailed reconstruction of root system architecture and its temporal dynamics under both controlled and semi-field conditions. Furthermore, root phenotyping is increasingly interpreted within an integrated whole-plant framework. The integration of organ-specific assessments with physiological phenomics leveraging spectral and thermal data enables the characterization of developmental plasticity and root-mediated processes, including water-use dynamics, nutrient acquisition, and canopy stress responses under heterogeneous field conditions. These approaches link root traits such as rooting depth and spatial distribution to canopy-level physiological responses under stress. Despite these advances, significant bottlenecks persist in data interoperability, analytical scalability, and protocol standardization. Future progress will require integration of root phenomics with genomics, predictive modelling, and digital twin frameworks to improve resource-use efficiency, yield stability, and climate resilience in global cropping systems.
Why it matches plant phenotyping methods根系フェノタイピングの高スループット画像化、計算ツール、機械学習統合、データ標準化を中心に扱う方法論レビューであり、植物形質の取得・解析手法が主題である。
titleAdvances in root phenotyping: high-throughput imaging, computational tools, and integrative approaches for crop improvement.
Forest digital twins play a crucial role in modern precision forestry by supporting biomass estimation and carbon cycle monitoring. However, existing 3D reconstruction methods struggle to simultaneously achieve metric-level structural accuracy and visual realism in complex understory environments. This study proposes a semantically constrained 3D Gaussian Splatting framework that fuses handheld LiDAR point clouds with unmanned aerial vehicle imagery. First, a multi-modal fusion mechanism is constructed to extract geometric anchors from registered LiDAR data for precise 3DGS spatial initialization, which mitigates rendering artifacts and geometric drift caused by poor initialization in purely visual methods. Second, a semantic regularization optimization strategy is proposed to realize differentiated modeling of tree trunks and canopies, effectively balancing the structural accuracy of rigid trunks and the photorealistic rendering of non-rigid canopies. Experiments conducted on three study plots demonstrate that the proposed approach achieves an average PSNR of 24.94 dB, SSIM of 0.773, and LPIPS of 0.231 across all plots, outperforming standard NeRF and baseline 3DGS, while enabling DBH estimation with R2 = 0.848 and RMSE = 2.705 cm. This method provides a solution for high-fidelity forest digital twin construction in open-canopy forest environments such as urban and campus forests.
Why it matches plant phenotyping methodsLiDAR・UAV画像を統合した3D再構成法を開発し、樹幹・樹冠の構造モデル化とDBH推定を評価しており、植物形質取得が中心的な技術貢献である。
abstractThis study proposes a semantically constrained 3D Gaussian Splatting framework that fuses handheld LiDAR point clouds with unmanned aerial vehicle imagery.
The rapid and nondestructive classification of maize kernels is of great significance for seed screening and quality evaluation. Existing hyperspectral image classification methods based on the Mamba architecture can effectively represent spectral and spatial features; however, they still face limitations in time-frequency analysis and multimodal feature fusion. In addition, traditional approaches often rely heavily on spectral preprocessing, which may introduce additional errors and compromise the model's robustness and generalization ability. To address these challenges, this paper proposes a novel cross-modal classification framework named CD-TriMamba, which jointly leverages hyperspectral data and visible-light images for comprehensive feature extraction and deep fusion. Specifically, an innovative feature extraction module is designed, consisting of a Spectral Curvelet Convolution (SCC) module for hyperspectral data and a Curvelet-Decomposed Convolution (CDC) module for spatial modeling. A feature rearrangement mechanism is further introduced to mine critical information from both spectral and spatial modalities. Finally, a ConvNeXt-guided tri-branch cross-fusion structure (TriMamba) is constructed to achieve deep collaboration and efficient integration between spectral and spatial features. Experimental results demonstrate that the proposed model achieves outstanding performance in seed classification, with an accuracy (Acc) of 99.2% and a Kappa value of 99.1%. These results strongly confirm the effectiveness and broad application potential of cross-modal feature fusion in maize kernel classification.
Why it matches plant phenotyping methodsマルチモーダル画像からトウモロコシ種子の健全性を推定する新規分類フレームワークを開発しており、種子状態の取得・抽出手法が研究の中心である。
abstractthis paper proposes a novel cross-modal classification framework named CD-TriMamba, which jointly leverages hyperspectral data and visible-light images for comprehensive feature extraction and deep fusion.
Field / plotMultimodalWhole plant / canopy / plot / fieldClassification
California agriculture faces the combined pressures of severe drought, high crop-waste rates, and unaffordable commercial precision-agriculture platforms, which together disproportionately affect small and mid-sized growers. This paper proposes an integrated planthealth monitoring platform that combines a Raspberry-Pi field node for image capture and environmental sensing, a multimodal vision model for species identification and health assessment, and a Flutter mobile client that presents results to the grower through a simple dashboard. The client implements a layered fallback between the live Pi, an on-disk cache, and a bundled sample dataset so that it remains functional under intermittent connectivity, and it caches plant images transparently to accelerate repeated views. Two experiments evaluated the system: species identification reached 87.5 percent accuracy across four visually similar species, and time-to-first-paint ranged from 0.38 seconds on cached data to 1.34 seconds under degraded networks. The platform demonstrates that practical precision agriculture is achievable at consumer-hardware scale.
Why it matches plant phenotyping methods植物画像を取得し、植物の健康状態を評価する統合モニタリング基盤を開発・評価しており、画像ベースの植物状態推定が中心です。
abstractThis paper proposes an integrated planthealth monitoring platform that combines a Raspberry-Pi field node for image capture and environmental sensing, a multimodal vision model for species identification and health assessment
Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationCountingGrowth / development / phenology
Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.
Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画データセットであり、自動フェノロジー検出と空間検証のためのベンチマーク基盤が中心である。
abstractThe Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions.
Reproduction assets foundThe paper describes a public multi-modal Actinidia chinensis phenology dataset (annotated images, georeferenced videos, ground-truth counts) deposited on Zenodo, plus authors' MIT-licensed preprocessing scripts on GitHub. CVAT and FiftyOne are generic third-party tools and excluded.Dataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available
at Zenodo: https://doi.org/10.5281/zenodo.17371025.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:12 lines:1-92Code · publicCustom scripts for dataset preparation are publicly available under the MIT License at https://github.com/Open asset ↗GitHubpdf-page:12 lines:1-92Code / dataset availability confirmedCrossref · Europe PMC · checked 15 Sept 2026
Abstract Plant diseases are a serious danger to the world’s food security, because they lower agricultural output and increase economic losses. Due to subjectivity, fluctuating lighting, and environmental unpredictability, traditional visual examination techniques are frequently incorrect. The Excess Green (ExG) vegetation index and pseudo-thermal representations produced from RGB pictures are two synthetically developed complementary representations that are integrated with RGB imagery in this study’s lightweight multimodal deep learning system to address these issues. Histogram shifting and pseudo-infrared color mapping are used in a reproducible picture alteration pipeline to create the pseudo-thermal modality, which allows for extra visual signals without the need for specific thermal sensors. In order to classify plant diseases while preserving computational efficiency, the suggested framework uses MobileNetV3-Small backbones to extract modality-specific characteristics. This is followed by feature-level fusion. The publicly accessible Ginger Leaf Dataset, which includes RGB pictures of ginger leaves in four different conditions—Damage-Pest, Dehydrated, Healthy, and Leaf-blight—was used for the experiments. For training, validation, and testing, the dataset was split using a stratified 70:15:15 split. Python-based preprocessing procedures were used to create the extra modalities (ExG and pseudo-thermal representations) from the original RGB images. The experimental results show that the combination of the representations with RGB images can enhance the classification performance compared with the unimodal RGB-based models. Ablation experiments are also conducted to examine the contributions of different modalities to the overall categorization accuracy. The experimental results show that plant disease recognition can be improved with the help of efficient computing by combining lightweight convolutional neural networks with computationally generated visual representations.
Why it matches plant phenotyping methodsRGB画像からExG・疑似熱画像を生成し、植物葉の病害状態を分類するマルチモーダル手法が研究の中心であり、アブレーション評価も実施している。
titleHybrid deep learning-based multimodal framework for plant leaf disease classification using RGB, Excess Green (ExG), and pseudo-thermal representations with MobileNetV2
Reproduction assets foundThe paper's phenotyping experiments use the publicly available Ginger Leaf Dataset (RGB leaf images of four ginger leaf conditions), with a public GitHub repository and dataset website. The authors' derived ExG/pseudo-thermal representations and preprocessing scripts are only available upon request, so they do not yetDataset · publicor multispectral images
IEEE Geosci. Remote Sens. Lett. 2025
10.1109/LGRS.2025.XXXXXXX
Ulku, I., Tanriover, O. O. & Akagündüz, E. Cross-band correlation-aware interactive fusion for multispectral images. IEEE Geosci. Remote Sens. Lett.
10.1109/LGRS.2025.XXXXXXX
(2025).
10. Wong, J. Ginger Leaf Dataset. GitHub Repository (2023). https://github.com/wongjay1941/Ginger-Leaf-Dataset
11.
Bhakta I
A novel plant disease prediction model based on thermal images using modified deep convolutional neural network
Precis. Agric. 2023 24 23 39
10.1007/s11119-022-09927-x
Bhakta, I. et al. A novel plant disease prediction model based on thermal images using modified deep convolutional neural network. Precis.Open asset ↗https://github.com/wongjay1941/Ginger-Leaf-Datasetlines:580-681Plant phenotyping relevance match · UnverifiedCrossref · checked 15 Sept 2026
Plant disease and pest surveillance is undergoing a profound technological transition. Conventional crop protection has historically depended on episodic field scouting, expert visual inspection, and broad-spectrum preventative spraying, all of which are constrained by labour intensity, uneven diagnostic accuracy, and weak temporal resolution. In contrast, recent advances in artificial intelligence, deep learning, the Internet of Things, remote sensing, and edge computing have enabled crop-health monitoring systems that are more continuous, data-rich, and spatially explicit. This review analyses the evolution of AI-driven plant disease and pest surveillance, with particular attention to how image-based deep learning, connected environmental sensing, unmanned aerial vehicle platforms, cloud-edge infrastructures, and multimodal analytics are reshaping next-generation crop protection. The article argues that the central innovation is not merely automated diagnosis, but the emergence of surveillance ecosystems capable of recognising symptoms, estimating risk, localising hotspots, and informing more selective intervention. The review synthesises major developments in convolutional neural networks, object detection, semantic segmentation, transfer learning, domain adaptation, transformer-based computer vision, anomaly detection, environmental time-series modelling, and multimodal analytics. It also evaluates the practical obstacles that still limit real-world deployment, including dataset bias, annotation uncertainty, poor cross-domain generalisation, limited interoperability, energy and connectivity constraints, weak model explainability, and uneven economic accessibility. The article further considers how AI-based surveillance may strengthen integrated pest management by supporting earlier warning, more precise treatment timing, reduced blanket pesticide use, and stronger alignment between biological risk and management action. It concludes that the future of crop protection will depend less on isolated improvements in benchmark accuracy and more on the development of trustworthy, scalable, and biologically meaningful surveillance systems that can support sustainable decisions under real agricultural conditions.
Why it matches plant phenotyping methods植物病害の症状認識・リスク推定・ホットスポット局在化を行う画像解析、センサー、UAV、マルチモーダル手法を中心にレビューしており、植物の病害状態を推定するフェノタイピング手法レビューに該当する。
abstractThis review analyses the evolution of AI-driven plant disease and pest surveillance, with particular attention to how image-based deep learning, connected environmental sensing, unmanned aerial vehicle platforms, cloud-edge infrastructures, and multimodal analytics are reshaping next-generation crop protection.
Rhizosphere oxidation is a key adaptive mechanism in reductive soil environments, in which oxygen released from roots alters rhizosphere redox conditions and regulates biogeochemical processes. Rice plants possess an internal oxygen transport system, and radial oxygen loss (ROL) from roots is closely associated with root development. However, the spatial patterns of ROL in soil and their relationships with root traits remain poorly characterized. In this study, we developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice. Daily time-course tracking of individual crown roots revealed dynamic changes in the spatial distribution and magnitude of rhizosphere oxygen in relation to root elongation and aging. Root thickness was positively correlated with dissolved oxygen levels near root tips. Genotypic comparisons further identified a cultivar with reduced rhizosphere oxidation despite possessing thicker roots among the tested genotypes, thereby indicating the involvement of additional physiological processes. Overall, these findings demonstrate that rhizosphere oxidation is regulated by root growth stage and thickness and dynamically modulated during root development.
Why it matches plant phenotyping methods平面酸素オプトードとX線CTを統合したマルチモーダル画像システムを開発し、イネ根の発達と根圏酸化を時系列・空間的に測定しているため、植物フェノタイピング手法が研究の中心である。
abstractwe developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice.
Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate detection of such diseases is crucial for improving both crop yield and quality. While numerous deep learning approaches rely solely on image data for disease identification, they often overlook the complementary value of textual information in enhancing visual analysis. To address this limitation and effectively fuse features from different modalities, we propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition. Our approach utilizes the Zhipu.ai multi-modal model to generate comprehensive textual descriptions of diseased crop leaves, including global description, local lesion description, and color-texture description. These textual descriptions are then encoded into feature embeddings, while visual features are extracted using the ShuffleNet-v2 model as the image encoder. Subsequently, a cross-attention module aligns and fuses the two modalities, and a gated fusion module enables dynamic feature selection during the fusion process. Extensive evaluations on the Soybean Disease and PlantVillage datasets demonstrate that our method outperforms existing image-based models in terms of accuracy. Specifically, our model achieves recognition accuracies of 99.04% and 99.12% on the respective datasets, surpassing the ShuffleNet-V2 model by 1.09% and 2.53%, respectively. These results highlight the effectiveness of Cross-Model learning in integrating visual and textual cues for accurate and efficient disease recognition, offering a scalable solution for crop disease diagnosis.
Why it matches plant phenotyping methods植物葉の病徴を画像と言語情報から認識する融合フレームワークを開発し、複数データセットで既存手法と比較評価しているため、植物フェノタイピング手法が中心である。
abstractwe propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition.
Reproduction assets foundThe paper's crop disease recognition experiments use two openly available image datasets, both with explicit public availability statements in the Data Availability section: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). No author analysis code, trained models, or generated text-annotaitDataset · publicThe datasets utilized in this study are openly accessible. The soybean dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3.Open asset ↗Dryad · 10.5061/dryad.41ns1rnj3html-lines:403-424Dataset · publicThe plantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-dataset.Open asset ↗Kaggle · plantvillage-datasethtml-lines:403-424Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Abstract Timely and accurate plant disease detection is important for enhancing agricultural productivity and promoting sustainability. The study introduces Multimodal Adaptive Fuzzy-based Deep Neural Network (MAF-DNN) for classification of plant diseases. The proposed method combines fuzzy logic with multimodal data fusion to effectively address the complex interactions and uncertainties in agricultural datasets. The MAF-DNN employs a robust adaptive fuzzy framework with dynamic rule optimization and integrates Hyperspectral Imaging Data (HID) with RGB imaging data to acquire detailed spectral information and high-resolution visual cues for disease classification. The multimodal fusion enhances the model’s ability to capture intricate patterns that relate to plant health, improving the accuracy of disease classification. The experimental results showed that the MAF-DNN outperforms traditional models by achieving an accuracy of 97.8%, precision of 96.5%, recall of 98.2%, and F1-score of 97.3%. Additionally, the adaptive design reduces computational overhead, increases efficiency, and improves scalability for large-scale agricultural applications. The MAF-DNN represents a significant advancement in plant disease classification and provides a robust and efficient solution for precision agriculture.
Why it matches plant phenotyping methods植物病徴を画像から分類するマルチモーダル画像・深層学習手法の開発と性能評価が中心であり、植物の病害状態を直接推定するため。
abstractThe study introduces Multimodal Adaptive Fuzzy-based Deep Neural Network (MAF-DNN) for classification of plant diseases.
Reproduction assets foundThe paper uses two public Kaggle plant disease image datasets (New Plant Diseases Dataset and CCMT Plant Disease Dataset) as its phenotyping inputs and states that the authors' custom MAF-DNN code is publicly available on GitHub, with all three URLs given in the article and matching allowed URLs.Code · publicThe custom code used to develop and evaluate the proposed Multimodal Adaptive Fuzzy Deep Neural Network (MAF-DNN) framework is publicly available at: https://github.com/skbsangeetha/MAF-DNN-Plant-disease-classificationOpen asset ↗skbsangeetha/MAF-DNN-Plant-disease-classificationhtml-lines:102-118Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · OpenAlex · checked 5 Sept 2026
Introduction With the continuous advancement of smart agriculture, multi-modal remote sensing based on unmanned aerial vehicles (UAVs) offers new technical approaches for monitoring and managing crop moisture in fields. However, significant challenges remain in developing high-precision field-scale crop Plant Moisture Content (PMC) prediction models and translating them into actionable irrigation strategies. Methods This study focuses on winter wheat, employing field experiments with PMC and water use efficiency (WUE) as indicators of crop water status. Vegetation indices (VIs) derived from UAV data were used to construct a leaf area index (LAI) inversion model. Crop Height was extracted from oblique photogrammetry point cloud data. By combining the Penman-Monteith equation with dual crop coefficients, an improved evapotranspiration (ET) model was developed, utilizing multispectral data from UAVs, thermal infrared data, point cloud-derived plant height, and LAI inversion results. Further utilizing VIs, temperature indices (TIs), and machine learning algorithms (Random Forest Regression (RFR), Back Propagation Neural Network (BPNN), Partial Least Squares Regression (PLSR), and Support Vector Regression (SVR), we established PMC prediction models for winter wheat at different growth stages. These models, integrated with WUE, form the basis for an irrigation scheduling optimization framework at the field scale. Results Results indicate that VIs, the difference between canopy temperature and air temperature (ΔT), Crop Water Stress Index (CWSI), and ET exhibit varying correlations with PMC during three critical growth stages of winter wheat, with ET showing the highest correlation during the jointing and heading stages (absolute correlation coefficient |r| ≥ 0.639). Compared to PMC prediction models constructed with different combinations of VIs, ET, VIs+ET, and VIs+TIs, the model employing the RFR algorithm with multimodal inputs (VIS+TIs+ET) demonstrated the best performance. The model’s predictive accuracy gradually improved across all growth stages, peaking during the grain-filling stage, with the coefficient of determination(R 2 ) of 0.900 and a normalized root mean square error (nRMSE) of 2.688%. Optimal WUE varied across growth stages under different irrigation treatments. The highest values were achieved at the jointing stage under treatment W3 (PMC = 81.8%), and at the heading and grain-filling stages under treatment W1 (PMC = 76.8% and 64.0%, respectively). Discussion The study suggests that stage-specific irrigation scheduling based on PMC thresholds can improve overall water use efficiency. This study shows that integrating multi-modal UAV data with machine learning and an improved ET model enables high-precision PMC monitoring, supporting data-driven irrigation scheduling in precision agriculture.
Why it matches plant phenotyping methodsUAVマルチモーダルデータと機械学習により、作物水分状態(PMC)、LAI、草高、蒸発散量を推定する手法を開発・評価しており、フェノタイピング手法が研究の中心である。
abstractCrop Height was extracted from oblique photogrammetry point cloud data.
Field / plotMultimodalLeafWhole plant / canopy / plot / fieldClassificationObject detectionStress / disease detectionDisease symptoms / severityWater status / transpiration
Crop losses caused by disease, water stress, and delayed field intervention remain a major challenge for small and medium farmers. Conventional advisory systems often depend on manual inspection or cloud-only diagnosis, which can be slow in rural environments where connectivity is limited. This paper proposes a multimodal edge-intelligence framework that combines leaf-image analysis, soil-moisture sensing, weather context, and lightweight decision rules to provide early crop disease detection and irrigation advisory. The system uses a compact convolutional neural network for visual symptoms and a sensor-fusion module for environmental risk estimation. By running inference near the field, the framework reduces latency and protects farm data while still supporting periodic cloud synchronization. Simulated evaluation shows 91.8% disease classification accuracy, 16.4% water saving, and faster advisory delivery compared with image-only and rule-based baselines.
Why it matches plant phenotyping methods葉画像から作物の病徴・病害状態を推定するCNNとセンサ融合手法の開発・評価が中心であり、植物状態の表現型計測に該当する。灌漑助言部分も含むが、病害検出手法が明示的に評価されている。
titleMultimodal Edge Intelligence for Crop Disease Detection and Irrigation Advisory in Precision Agriculture
Unmanned aerial vehicles (UAVs) are broadly used for high-throughput plant phenotyping, yet their long-term use in public-sector research is increasingly challenged by regulatory restrictions and reliance on proprietary platforms. This study presented a regulation-compliant, modular multi-sensor unmanned aerial system (UAS) designed to deliver flexible, high-quality phenotyping data without dependence on restricted ecosystems. A dual-mount, open-architecture payload integrated RGB, multispectral, and thermal sensors, enabling simultaneous acquisition of structural, spectral, and thermal information within a unified workflow. Field validation in a lantana (Lantana camara) breeding trial demonstrated high-precision multi-sensor data fusion and reliable trait extraction. Spatial co-registration achieved centimeter-level accuracy, with alignment errors of 0.88 cm (multispectral) and 3.23 cm (thermal) relative to the RGB reference. UAV-derived canopy height closely matched ground measurements (R2 up to 0.98; RMSE as low as 1.57 cm), while canopy coverage estimates showed consistency across sensing modalities (R2 = 0.99; RMSE = 0.02 m2). Calibrated thermal orthomosaics provided robust canopy temperature estimation (RMSE = 3.13 °C), supporting a quantitative assessment of plant physiological status. Together, these results demonstrate that a regulation-compliant, open-architecture UAV platform can achieve high accuracy in multi-modal phenotyping while maintaining flexibility and cost efficiency. This work demonstrates a scalable and sustainable framework for UAV-based phenotyping, enabling researchers to adapt to evolving regulations while advancing data-driven crop improvement.
Why it matches plant phenotyping methods植物フェノタイピング用のマルチセンサーUAVプラットフォームを設計・検証し、植物形質の抽出精度を評価しているため、方法が中心的である。
abstractThis study presented a regulation-compliant, modular multi-sensor unmanned aerial system (UAS) designed to deliver flexible, high-quality phenotyping data without dependence on restricted ecosystems.
ABSTRACT - Sugarcane is one of the most important commercial crops worldwide but its productivity is greatly affected by diseases such as red rot, rust, mosaic, smut and yellow leaf disease. Conventional disease detection techniques are based on manual inspection which is a time-consuming, labor-intensive and error prone process. This paper gives a detailed review of the deep learning methods for the automated detection of sugarcane diseases with a special focus on the fusion methods of stem and leaf features. Different deep learning architectures such as CNN, VGG, ResNet, EfficientNet, DenseNet, MobileNet, and YOLO are analyzed and compared in terms of accuracy, efficiency, and deployment capability. The study also explores multimodal approaches, such as hyperspectral imaging, thermal imaging and environmental data integration, to enhance prediction performance. Reported results show that advanced models like EfficientNet-B7 and DenseNet201 achieve accuracies above 99%, while lightweight models like MobileNet allow for real-time mobile deployment. The review highlights significant research gaps such as small datasets, lack of stem-leaf fusion studies, no severity classification, and real-world deployment issues. Future research directions are related to explainable AI, multimodal fusion, lightweight edge computing models, and precision agriculture applications for sustainable sugarcane cultivation. Key Words: Sugarcane disease detection, Deep learning, CNN, Stem-leaf fusion, Computer vision, Precision agriculture.
Why it matches plant phenotyping methodsサトウキビ病害の画像・深層学習による検出手法を中心にレビューしており、植物の病態を観測・推定するフェノタイピング手法レビューに該当する。
abstractThis paper gives a detailed review of the deep learning methods for the automated detection of sugarcane diseases with a special focus on the fusion methods of stem and leaf features.
This study proposes the Hydroponic Plant Growth Analysis System (HPGAS), a public-data-based preliminary framework for multimodal plant growth state analysis toward future filter-free aquaponic validation. The HPGAS integrates plant images, water quality signals, and environmental signals to estimate an image-centered growth index, growth stage, and proxy abnormal state probability. Because no public dataset jointly provides plant images, direct growth labels, fish metabolic variables, suspended solids, and nitrification-related measurements from a real filter-free aquaponic system, this study is not a direct operational validation. A two-stage evaluation was conducted using the Autonomous Greenhouse Challenge (AGC), HydroGrowNet, and two aquaponic Internet of Things (IoT) water quality datasets. Stage 1 implemented dataset loaders, image–sensor alignment, proxy label generation, and unimodal and fusion baselines. Stage 2 expanded handcrafted image and sensor-context features and adopted month-wise hold-out evaluation. The image-only model achieved the best growth index regression performance, with a root mean square error (RMSE) of 0.0492 ± 0.0187, whereas the fusion model showed a RMSE of 0.0837 ± 0.0196. Conversely, the fusion model achieved the best proxy abnormal state classification performance, with a F1 score of 0.9695 ± 0.0057 under the clean condition, decreasing to 0.9232 ± 0.0263 under sensor dropout and 0.9132 ± 0.0169 under image noise. Under sensor dropout, the fusion model was more stable than the sensor-only model, whereas under image noise it degraded more than the image-only model. These results indicate that multimodal fusion is most useful for proxy abnormal state classification and robust state interpretation, rather than universally superior scalar growth regression. The HPGAS provides a reproducible baseline for future real filter-free aquaponic experiments, while its operational validity remains to be tested using real filter-free aquaponic data.
Why it matches plant phenotyping methods植物画像とセンサーデータを統合し、成長指数・成長段階・異常状態確率を推定する再現可能な解析フレームワークを開発・評価しており、植物表現型の取得・推定手法が中心である。
abstractThis study proposes the Hydroponic Plant Growth Analysis System (HPGAS), a public-data-based preliminary framework for multimodal plant growth state analysis
Abstract To address the problem of fine branch identification and pruning decision for dormant apple trees, this study proposes a 3D point cloud branch recognition method integrating Neural Radiance Fields (NeRF) and the PointNeXt network. This method employs the neural radiance field theory to construct a point cloud model of apple trees, achieving fine detail representation and providing a high-precision, high-standard dataset for subsequent branch pruning experiments. First, a panoramic video is captured by circling the fruit tree, and a multi-view image sequence is obtained through frame sampling. Subsequently, the Structure from Motion (SfM) algorithm is employed for sparse reconstruction to recover the pose information of the images. On this basis, a neural radiance field model is trained. Hierarchical sampling is performed using ray casting, and the sampled points, combined with positional encoding, are fed into a multi-layer perceptron (MLP). The radiance field is then generated via volume rendering, from which a high-fidelity 3D point cloud model of the fruit tree is derived. Finally, the point cloud is processed using the PointNeXt semantic segmentation network to achieve the identification and segmentation of branches to be pruned and branches to be retained. To verify the effectiveness of the method, this study reconstructed point cloud models of dormant apple trees and selected 10 of them for experimental analysis. The algorithm achieved an average overall recognition accuracy of 75.15% and an average false negative rate (FNR) of 24.85%. The experimental results demonstrate that the proposed method constructs a 3D point cloud model with multi-scale, multi-modal, and high-precision phenotypic information at a relatively low cost. It not only overcomes the limitations of traditional 3D reconstruction methods, such as insufficient point cloud accuracy and difficulty in accurately identifying thin branches, but also effectively mitigates the high misrecognition rate observed in conventional branch recognition approaches. This provides technical support for unmanned agricultural machinery pruning in orchards and holds significant implications for achieving precision agriculture and sustainable development.
Why it matches plant phenotyping methodsNeRFとPointNeXtを用いてリンゴ樹の3D点群を構築し、剪定対象枝を認識・分割する手法が研究の中心であり、植物の形態・構造状態を直接推定して性能評価している。
abstractthis study proposes a 3D point cloud branch recognition method integrating Neural Radiance Fields (NeRF) and the PointNeXt network.
Apple moldy core disease is a major pathogenic disease that severely degrades the postharvest quality of apples, and its early internal lesions cannot be directly identified through visual appearance observation. To realize efficient and non-destructive early diagnosis, this study proposes a multimodal image coding method fusing Visible-Near Infrared Spectroscopy (Vis-NIR) and Electronic Nose (E-nose) data, which is combined with the SE-ResNet18 deep learning model for disease classification. By virtue of coding techniques including Gramian Angular Field (GAF), Markov Transition Field (MTF), and Recurrence Plot (RP), one-dimensional time-series and spectral data were converted into image representations, so as to visualize their spatiotemporal patterns and enhance subtle disease-related features. On this basis, a two-branch SE-ResNet18 model based on the channel attention mechanism was constructed to improve the feature representation capability and achieve effective modal fusion. Experimental results show that the multimodal fusion model achieves a classification accuracy of 95.93%, which is significantly superior to single-modal methods, thus verifying the effectiveness of multi-source information complementarity. Ablation experiments further indicate that the SE attention module plays a crucial role in feature calibration and modal balance. Information entropy analysis reveals that the proposed method effectively enhances the discriminability of information during the feature extraction process. This study provides a solution with a clear theoretical basis and reliable performance for the non-destructive detection of early diseases in agricultural products, which has favorable application prospects and popularization potential.
Why it matches plant phenotyping methodsリンゴ果実の内部病変という植物器官の病態を対象に、Vis-NIR・E-noseのマルチモーダルセンシングと深層学習による非破壊分類法を開発・評価しており、表現型取得手法が中心である。
abstractTo realize efficient and non-destructive early diagnosis, this study proposes a multimodal image coding method fusing Visible-Near Infrared Spectroscopy (Vis-NIR) and Electronic Nose (E-nose) data, which is combined with the SE-ResNet18 deep learning model for disease classification.
Artificial intelligence (AI)-enabled camera sensor systems are increasingly transforming precision agriculture by providing non-destructive, rapid, and scalable methods for monitoring crop health. Two of the most critical applications are the detection of crop water stress and the assessment of pesticide requirement through pest, disease, and symptom recognition. This literature review synthesizes published work on RGB, thermal, multispectral, and hyperspectral imaging integrated with machine learning and deep learning methods for agricultural decision support. The reviewed studies show that thermal and hyperspectral imaging are particularly effective for water stress detection, whereas RGB and multispectral systems are highly practical for identifying disease symptoms, pest infestation, and spray targets. The literature further indicates a shift from simple classification toward real-time decision support, multimodal fusion, explainable AI, and precision input application. This review discusses core sensing technologies, major algorithmic approaches, research findings from key studies, present limitations, and future research directions. Overall, AI camera sensor systems offer substantial potential for reducing water wastage, minimizing excessive pesticide use, and improving sustainable agricultural productivity.
Why it matches plant phenotyping methods作物の水ストレスや病害症状を画像・センサーから推定する手法を中心に整理したレビューであり、植物状態の取得・推定方法が中核です。
abstractThis literature review synthesizes published work on RGB, thermal, multispectral, and hyperspectral imaging integrated with machine learning and deep learning methods for agricultural decision support.
Introduction Addressing the core bottleneck in traditional crop models-the disconnect between morphology and physiological function at the organ scale and their limited dynamic response to environmental changes-this study aimed to construct a multi-source data fusion maize growth model for simultaneous organ-scale simulation. Methods We developed a closed-loop Environment-Driven-Functional Response-Morphological Feedback (EDFM) architecture. By integrating environmental time-series data, RGB images, and 3D point clouds, we created a multimodal fusion model based on a gated attention network. This approach adaptively weights multi-source features and pioneers a bidirectional morphology-physiology feedback loop based on physiological development time (PDT) and NURBS surfaces. The WOFOST moisture response function was also improved. Results The model significantly enhanced the simulation accuracy of organ-scale growth, reducing the root mean square error (RMSE) for plant height by 74.6% through a morphology-physiology dynamic weighting mechanism. More fundamentally, it resolved the disconnect between morphological and physiological processes. The improved plant height prediction validates the model's effectiveness at the organ scale. Discussion The pioneering "physiology-morphology" parallel simulation architecture provides an interpretable theoretical model and robust quantitative tools for designing high-photosynthetic-efficiency plant architecture and enabling precision water-fertilizer management.
Why it matches plant phenotyping methodsRGB画像・3D点群・環境データを統合し、器官スケールの形態と生長をシミュレーションする手法を開発しており、植物形質(草丈など)の推定が中心的な技術貢献である。
abstractWe developed a closed-loop Environment-Driven-Functional Response-Morphological Feedback (EDFM) architecture.
The intelligent transformation of agriculture places plant growth prediction as a critical component for ensuring food security, optimizing resource allocation, and enhancing sustainable productivity. Traditional methods reliant on empirical or simplified mechanistic models struggle with the nonlinearity, high dimensionality, and spatiotemporal heterogeneity inherent in agro-ecological systems. This study investigates the paradigm shift enabled by agricultural big data integrating multi-source, real-time streams from IoT sensors, satellites, UAVs, and farm management systems. We propose a ``Multi-source Data Assimilation and Hybrid Intelligence'' (MDA-HI) framework that synergistically couples process-based crop models with ensemble machine learning algorithms---including Transformer-based architectures and Physics-Informed Neural Networks---within a holistic pipeline encompassing multi-modal data fusion, hybrid modeling, and scalable deployment. Empirical validation across major crops (rice, wheat, maize, tomato) in diverse eco-regions of China (2023--2025) demonstrates significant improvements: the MDA-HI model achieved average RMSE reductions of 42.7% for yield prediction and 38.1% for key phenological stage prediction relative to best-in-class standalone models. A large-scale case study on rice-wheat rotation systems showed that data-driven prescriptions reduced nitrogen fertilizer use by 22.5% and irrigation water by 18.3% while increasing yield by 5.1%. The study further establishes a five-dimensional evaluation system covering accuracy, robustness, interpretability, scalability, and economic benefit. Remaining challenges include edge computing for real-time inference, federated learning for privacy-preserving collaboration, and explainability of complex ``black-box'' models. This research concludes that agricultural big data constitutes a foundational catalyst for predictive, precise, and proactive cognitive agriculture, with profound implications for global food system resilience.
Why it matches plant phenotyping methods農業ビッグデータを用いて生育・収量・フェノロジーを推定するMDA-HI手法を提案し、複数作物・地域で性能検証しており、植物形質推定手法が研究の中心である。
abstractWe propose a ``Multi-source Data Assimilation and Hybrid Intelligence'' (MDA-HI) framework that synergistically couples process-based crop models with ensemble machine learning algorithms---including Transformer-based architectures and Physics-Informed Neural Networks---within a holistic pipeline encompassing multi-modal data fusion, hybrid modeling, and scalable deployment.
Early-stage frost damage in citrus fruits is difficult to detect because external symptoms are often weak or absent, hindering intelligent robotic sorting in postharvest scenarios. To address this challenge, this study proposes a robotic multimodal tactile sensing approach inspired by human mechanoreception for frost-damage detection during grasping. A robotic gripper equipped with a 6×6 pressure matrix sensor and a piezoelectric vibration sensor was used to capture complementary tactile cues during standardized fruit handling, enabling the perception of subtle mechanical changes associated with early frost injury. Using 240 Citrus reticulata 'Hong Mei Ren' fruits under controlled experimental conditions, a Transformer-based multimodal fusion network was developed to jointly model pressure and vibration sequences for binary classification of normal and frost-damaged fruits. Across repeated stratified random-split experiments, the proposed method achieved a mean classification accuracy of 93.1%. Comparative experiments showed that the fusion model outperformed representative sequence-learning baselines, and ablation analysis confirmed that pressure-vibration fusion was more effective than either single modality alone. Attention-based temporal attribution further revealed that the most informative cues were concentrated in the initial contact and early loading stages, indicating the importance of early transient mechanical responses for frost-damage discrimination. Overall, the proposed approach demonstrates the feasibility of grasp-based robotic frost-damage detection under controlled experimental conditions.
Why it matches plant phenotyping methods柑橘果実の凍害状態を圧力・振動センサーで取得し、マルチモーダル融合により分類する手法の開発が中心であり、単なる生物学的実験の測定ではない。
abstractthis study proposes a robotic multimodal tactile sensing approach inspired by human mechanoreception for frost-damage detection during grasping.
The leaf area index (LAI) is a key parameter for characterizing crop growth and water use efficiency. Therefore, efficient and accurate monitoring of LAI is essential for precision rice management. To overcome the limitations of traditional LAI measurement methods, which are time consuming, labor intensive, and difficult to scale, this study proposes an inversion framework that integrates multi-source UAV remote sensing features with machine learning models. The framework incorporates color indices (CIs) derived from RGB imagery, vegetation indices (VIs) derived from multispectral data, texture features (TIs), and texture feature indices (TFIs), and employs six machine learning algorithms to develop optimized LAI estimation models for the rice booting stage. The results indicate that at a flight altitude of 30 m, the CNN model integrating CIs and TIs achieved an accuracy of R 2 = 0.815. At 60 m, the RF model combining VIs and TFIs showed superior performance, with an R 2 of 0.866. Further integration of CIs, VIs, and TFIs at 30 m produced the best results, increasing R 2 to 0.901, reducing RMSE to 0.273, and raising RPD to above 3.0. These findings demonstrate that TFIs significantly enhance the spectral-spatial representation capability of multispectral data, thereby improving model accuracy. The combined use of CIs and VIs across different sensors compensates for the inherent limitations between spectral and spatial information, while the integration of multi-resolution TIs and TFIs effectively overcomes the constraints of single-source data. Overall, the proposed approach provides a robust and efficient solution for high-precision LAI estimation during critical growth stages of rice, offering strong support for precision agricultural management.
Why it matches plant phenotyping methodsUAV画像・マルチスペクトル特徴量と機械学習によるイネLAI推定フレームワークの開発・性能評価が研究の中心であり、植物形質の取得手法に該当する。
abstractthis study proposes an inversion framework that integrates multi-source UAV remote sensing features with machine learning models.
Disease prevention and water management are important to all the crops, particularly rice and sugarcane production in India. The article proposes a reinforcement learning (RL) based intelligent irrigation management system that is capable of optimising water consumption and crop nutrition in response to the changing agricultural climatic conditions. Decentralised reinforcement learning (RL) is used in a network of irrigation agents that utilise soil and microclimate sensor networks to set the terms of water allocation, water use efficiency (WUE) and crop health. At the same time, deep convolutional networks can be used to differentiate between plant stress/disease and leaf images and take applicable proactive actions. It is a framework that incorporates satellite-derived indices (NDVI, EVI, land surface temperature) with local sensor measurements and image-based health measurements through multimodal deep learning. Far-reaching simulations (including Indian climate and crop calendars) demonstrate that the multi-agent system lowers water consumption and preserves the yields and properly notifies stressed plants. The scores of disease detection with plantvillage-based fine-tuned on rice (120 (3 disease types) and 3829 (5 disease types) and sugarcane (2569 images for all disease types, Convolutional Neural Network (CNN) yield results of >98 % accuracy. Crop mapping (rice/sugarcane) Satellite/LSTM-based crop mapping (with Sentinel-1 / Sentinel-2) achieves more than 97 % accuracy. The suggested structure provides a data-driven, scalable system for precision agriculture to enhance the management of irrigation periods and crop health. Simulation experiments show that the RL-based controller can reduce water consumption while preserving optimal soil moisture levels when compared to rule-based irrigation strategies.
Why it matches plant phenotyping methods画像・衛星・センサーを統合して植物ストレス/病害状態を推定するマルチモーダル基盤が提案され、病害検出性能も評価されているため、植物表現型推定が実質的な構成要素である。
abstractdeep convolutional networks can be used to differentiate between plant stress/disease and leaf images
Reproduction assets foundThe paper reports simulation-based experiments using public leaf-image datasets. The only paper-specific public asset explicitly identified is the Kaggle rice leaf diseases dataset (vbookshelf/rice-leaf-diseases) cited as a data source for the rice disease fine-tuning set. No authors' code, trained models, or data dépDataset · publicConflict of interest: Authors do not have any conflict of interest 2026 Mar 31). Available from: https://www.kaggle.com/datasets/Open asset ↗Kagglepdf-page:16 lines:1-58Plant phenotyping relevance match · UnverifiedEurope PMC · checked 15 Sept 2026
Common beanMultimodalClassificationStress / disease detectionDisease symptoms / severity
Abstract Recent multimodal agricultural research emphasizes large-scale vision-language systems, while lightweight reproducible approaches remain underexplored for constrained deployments. We present AgroMM-GSF++, a compact image-text fusion framework for plant disease recognition. The model combines a small visual backbone and text-prototype branch with confidence-adaptive gating. To strengthen empirical evidence, we run repeated multi-seed stratified cross-validation (three seeds, repeated folds) and include stronger pretrained transfer baselines (EfficientNet-B0 and ViT-B/16 fine-tunes). We further evaluate on an external PlantVillage subset to test cross-dataset robustness under the same protocol family. On beans, transfer baselines lead absolute macro-F1 (EfficientNet-B0-FT: 0.6355±0.0340), while AgroMM-GSF++ reaches 0.5891±0.0552 with much lower latency than transfer models (9.80 vs 205.66 ms/image for EfficientNet-B0-FT). On the external subset, AgroMM-GSF++ improves over fixed lightweight fusion (0.8719 vs 0.8608 macro-F1) while remaining below transfer-heavy baselines. We provide an end-to-end reproducible workspace including scripts, datasets, figures, and camera-ready tables.
Why it matches plant phenotyping methods植物病害状態を画像から認識する軽量マルチモーダル手法を開発し、交差検証と外部ベンチマークで評価しており、植物フェノタイピング手法が研究の中心です。
abstractWe present AgroMM-GSF++, a compact image-text fusion framework for plant disease recognition.
Plant stress monitoring is invaluable in realizing sustainable agriculture because it enables the people practicing it to take early measures to counteract losses in yield caused by environmental stressors like drought and nutrient deficiencies, as well as caused by pathogen infections. The proposed study presents a new Multi-Modal Vision Transformer (MMViT) architecture that is designed to combine both thermal and RGB imagery to take detection to the next level. In order to support this methodology and further studies, we are now publicly releasing a new collection of synchronized thermo-RGB image pairs of stressed and healthy plants, collected both in controlled settings and in the field. The data is labeled to differentiate various stress phenotype and contains over 4286 of images, and hence forms a substantial platform to evaluate multimodal plant phenotyping methods. Empirical evaluations indicate that MMViT model achieves a general classification of 94.3% when using the two modalities, which is better than the single-modality ViT used on the thermal images (85.5%) and the RGB images (93.3%). These experimental results emphasize the performance of multimodal fusion whereby the other spectral cues are used to complement a stress classification. The described framework, together with the useful dataset, will contribute to the advancement of precision agriculture as it is an open and data-driven instrument to monitor plant health automatically.
Why it matches plant phenotyping methods熱画像とRGB画像を統合して植物ストレス表現型を分類するモデルを開発し、公開データセットと性能評価も提示しており、植物フェノタイピング手法が中心である。
abstractThe proposed study presents a new Multi-Modal Vision Transformer (MMViT) architecture that is designed to combine both thermal and RGB imagery to take detection to the next level.
Plant disease diagnosis in real-world agricultural environments is challenged by data scarcity, domain shift, privacy constraints, and limited edge-device resources. This paper proposes FMEL-FSDA, a Federated Multimodal Edge Learning framework with Few-Shot Domain Adaptation for robust field-based plant disease recognition. The framework integrates attention-based RGB–text feature fusion, privacy-preserving federated learning, rapid few-shot personalization, and uncertainty-aware inference within an edge-efficient architecture. Federated training enables collaborative learning across distributed farms without sharing raw data, while few-shot adaptation allows fast deployment to new regions using only 1–10 labeled samples per class. Experiments on the PlantWild in-the-wild dataset show that FMEL-FSDA outperforms centralized, federated, and few-shot baselines, achieving 93.78% accuracy, 93.33% F1-score, and 0.97 AUC. The model maintains strong performance under privacy mechanisms such as gradient perturbation and secure aggregation, reduces communication overhead by up to 4×, and supports low-latency edge inference. Uncertainty estimation and Grad-CAM-based explainability further enhance reliability by identifying low-confidence cases and highlighting disease-relevant regions. Overall, FMEL-FSDA offers a scalable, privacy-aware, and field-ready solution for intelligent plant disease diagnosis in precision agriculture.
Why it matches plant phenotyping methods植物病害の画像ベース認識手法を開発し、複数ベースラインとの比較評価と実フィールド性能検証を行っており、病害状態の取得・推定が研究の中心である。
abstractThis paper proposes FMEL-FSDA, a Federated Multimodal Edge Learning framework with Few-Shot Domain Adaptation for robust field-based plant disease recognition.
Robotic phytotronic systems in the agroindustrial complex ensure a reduction in the energy intensity of technological processes of intensive plant cultivation. Automated control of phytotrons using neural network computer vision provides early detection of abnormalities in plant development, contributing to the elimination of diseases and thermotherapy.
Why it matches plant phenotyping methods植物の発育異常・病気をニューラルネットワークによるコンピュータビジョンで早期検出する自動化フェノタイピングシステムが中心であり、植物状態の画像ベース推定に該当する。
abstractAutomated control of phytotrons using neural network computer vision provides early detection of abnormalities in plant development
Accurate identification of maize diseases is crucial for safeguarding global food security. Traditional image-based methods often struggle with lighting variations, occlusions, and noise, limiting their robustness and generalisation. Multimodal approaches that integrate visual and textual information have shown promise. However, these methods frequently require manually curated textual descriptions for each image, increasing data collection costs and limiting scalability and practical implementation. To address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA). This approach enforces category-level alignment between image and text modalities, enabling more accurate and interpretable cross-modal mapping. First, we construct cross-modal representations by aligning image and text modalities at the category level within a shared embedding space. Second, inspired by contrastive learning, we introduce a Cross-Modal Category Alignment (CMCA) loss based on category-level textual descriptions, reducing annotation complexity. Finally, we present an Efficient Channel-Spatial Hybrid Attention (CSHA) module that preserves inter-class boundaries while incurring minimal computational overhead, thereby enhancing feature discriminability under complex conditions. Experimental results on the maize subset of the PlantVillage dataset (MPVD) show that mIT-CMCA achieves 99.48% accuracy, 99.28% precision, 99.54% recall, and 99.41% F1-score. These results represent improvements of 0.24%, 0.13%, 0.17%, and 0.15% over the strongest vision-only baseline, MaxViT_tiny. On the self-built Maize Leaf-Field dataset (MLFD), the model achieves 93.67% accuracy, 93.76% precision, 93.67% recall, and 93.71% F1-score. It uses only 8.27 million parameters, which is 71.6% fewer than MaxViT_tiny. Its model size is 32.13 MB, which is 72.3% smaller. The proposed method also outperforms comparative models in robustness experiments under artificially added perturbations. These results demonstrate that mIT-CMCA achieves a favorable balance between accuracy and efficiency, making it suitable for practical agricultural deployment.
Why it matches plant phenotyping methodsトウモロコシ葉画像から病害状態を推定する画像・マルチモーダル手法の開発と性能評価が研究の中心であり、植物病害フェノタイピングに該当する。
abstractwe proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA).
Reproduction assets foundThe paper's maize disease identification analysis code and trained models are explicitly stated as publicly available in a GitHub repository. The phenotype image datasets (MPVD subset and self-built MLFD) are not publicly available and require contacting the corresponding author.Code · publicCode availability
The code and models are available in the GitHub repository at https://github.com/TANGFEILONG626/mIT-CMCA..Open asset ↗TANGFEILONG626/mIT-CMCAhtml-lines:673-695Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
This paper integrated multimodal remote sensing (RS) data with deep learning to develop a maize growth analysis and income prediction model based on CNN (Convolutional Neural Network)-Attention and Bi-LSTM (Bidirectional Long Short-Term Memory). Utilsing Landsat 8, Sentinel-2, and Sentinel-1 satellite data, the CNN-Attention network extracts vegetation features such as NDVI (Normalised Difference Vegetation Index), EVI (Enhanced Vegetation Index), and canopy density for corn growth stage identification and prediction. The study combined these features with meteorological, soil, and market data and fed them into a Bi-LSTM model for time series analysis to forecast corn income. The method achieved 98.3% accuracy in growth stage classification, with an RMSE (Root Mean Squared Error) of 0.04 and R² of 0.94 for canopy coverage prediction under 10-fold cross-validation, and an RMSE of 0.13 and R² of 0.94 for income prediction, showing high stability over time. Overall, this research supports data-driven precision agriculture through real-time monitoring and reliable economic forecasting.
Why it matches plant phenotyping methods衛星リモートセンシングと深層学習による作物生育段階・キャノピー密度の推定手法が中心的に開発・検証されており、植物状態の定量化を含むため。収益予測も扱うが、植物フェノタイプ推定部分が明示されている。
abstractintegrated multimodal remote sensing (RS) data with deep learning to develop a maize growth analysis and income prediction model
Pteris vittata, an arsenic-hyperaccumulating fern, is widely employed for phytoremediation of arsenic (As). Rapid, accurate assessment of As in P. vittata is crucial for evaluating its accumulation ability. In this study, P. vittata was analyzed using a spectral fusion of laser-induced breakdown spectroscopy (LIBS) and X-ray fluorescence (XRF). A total of 60 biological samples (roots and fronds) were collected and prepared as 180 compressed tablets for spectroscopic analysis, covering an As concentration range of 88-1956 mg kg -1 . Multivariate analysis methods were employed for full spectra and feature spectra, including partial least squares regression (PLSR), least squares support vector machine (LSSVM), extreme learning machine (ELM), random forest (RF), and adaptive weighting normalization-linear weighted network (AWN-LWNet). The best single-modality model, an XRF-based PLSR model built upon feature spectra selected by the Competitive Adaptive Reweighting Sampling (CARS) algorithm, achieved a prediction performance of R 2 P = 0.969, RMSE P = 54.13 mg kg -1 , and MAE P = 43.00 mg kg -1 . Spectral fusion further enhanced prediction accuracy, with high-level fusion outperforming low- and mid-level methods. The proposed feature-spectra-based decision fusion model achieved the best performance (R 2 P = 0.980, RMSE P = 43.69 mg kg -1 , MAE P = 32.49 mg kg -1 ), corresponding to reductions of 19.3% in RMSE P and 24.4% in MAE P compared to the best single-modality model. These results demonstrate that spectral fusion effectively integrates complementary information, improving the accuracy of As quantification in complex plant matrices. The proposed approach provides a rapid and non-destructive strategy for monitoring arsenic accumulation in phytoremediation plants.
Why it matches plant phenotyping methodsLIBS・XRFのスペクトル融合と機械学習により、植物組織中のヒ素蓄積量を非破壊推定する手法を開発・性能評価しており、植物状態の取得が中心です。
abstractThe proposed approach provides a rapid and non-destructive strategy for monitoring arsenic accumulation in phytoremediation plants.
Seeds are complex living systems that display rich diversity in morphology, physiology, biochemistry, and genetics. Yet phenotyping during seed dormancy remains hampered by limited imaging modalities, insufficient data integration, and underpowered intelligent analytics—constraints that impede the efficiency and accuracy of precision breeding and germplasm evaluation. In the AI-for-Science era, seed phenomics research urgently needs to establish an end-to-end “measure–compute–understand–apply” pipeline, spanning cross-scale multimodal data acquisition to insight. This article systematically reviews the technical evolution of dormancy-state seed phenotyping and delineates five stages—manual observation phenotypes, biochemical phenotypes, image-based phenotypes, digital seeds, and intelligent seeds—summarizing the defining features and principal limitations of each. The deep integration of advanced imaging with artificial intelligence offers new opportunities to overcome existing bottlenecks. Seed Imaging Omics has emerged to meet this need: leveraging multiscale, multidimensional imaging for comprehensive observation and multimodal data capture; coupling these data with multimodal fusion, foundation-model analysis, and virtual seed reconstruction to enable precise feature extraction and pattern discovery from large image corpora. These capabilities clarify complex traits, reveal morphology–function relationships, and advance systems-level understanding of seed biology, ultimately supporting precise germplasm management and evaluation, data-driven elucidation of biological mechanisms, and accelerated innovation in crop improvement. Looking ahead, continued progress in sensing and imaging, foundation models, and large-scale analytics will drive seed phenotyping toward “intelligent” systems capable of autonomous sensing, real-time analysis, and decision-making across the seed life cycle—transforming seeds from passive carriers of genetic and phenotypic information into smart units that integrate phenotypic logging, state monitoring, performance assessment, and management feedback.
Why it matches plant phenotyping methods種子休眠状態の表現型計測技術を体系的にレビューし、画像取得、マルチモーダル統合、特徴抽出、AI解析を中心に扱うため、植物フェノタイピング手法レビューとして含める。
abstractThis article systematically reviews the technical evolution of dormancy-state seed phenotyping
Unmanned aerial vehicles (UAVs) have become indispensable tools in precision agriculture and plant phenotyping, enabling the rapid, non-destructive assessment of crop traits across space and time. Equipped with RGB, multispectral, thermal, and other sensors, UAVs provide detailed information on canopy structure, physiology, and stress responses that can guide management decisions and accelerate breeding programs. Despite these advances, the downstream processing of UAV imagery remains technically demanding. Converting orthomosaics into standardized, biologically meaningful data often requires a combination of photogrammetry, geospatial analysis, and custom scripting, which can limit reproducibility and accessibility across research groups. We present drone2report, an open-source python-based software that processes orthomosaics from UAV flights to generate vegetation indices, summary statistics, derived subimages, and text (html) reports, supporting both research and applied crop breeding needs. Alongside the basic structure and functioning of drone2report, we also present five case studies that illustrate practical applications common in UAV-/drone-phenotyping of plants: (i) thresholding to remove background noise and highlight regions of interest; (ii) monitoring plant phenotypes over time; (iii) extracting information on plant height to detect events like lodging or the falling over of spikes; (iv) integrating multiple sensors (cameras) to construct and optimize new synthetic indices; (v) integrate a trained deep learning network to implement a classification task. These examples demonstrate the tool’s ability to automate analysis, integrate heterogeneous data and models, and support reproducible computation of agronomically relevant traits. drone2report streamlines orthorectified UAV-image processing for precision agriculture by linking orthomosaics to standardized, plot-level outputs. Its modular, configuration-driven design allows transparent workflows, easy customization, and integration of multiple sensors within a unified analytical framework. By facilitating reproducible, multi-modal image analysis, drone2report lowers technical barriers to UAV-based phenotyping and opens the way to robust, data-driven crop monitoring and breeding applications.
Why it matches plant phenotyping methods植物表現型取得のためのUAV画像処理ソフトウェアを開発し、植物高・倒伏などの形質抽出、マルチセンサー統合、再現可能な解析ワークフローを中心的に提示している。
abstractWe present drone2report, an open-source python-based software that processes orthomosaics from UAV flights to generate vegetation indices, summary statistics, derived subimages, and text (html) reports
Reproduction assets foundThe paper explicitly states that the code and data to reproduce its five case studies (thresholding, temporal vegetation indices, height analysis, multi-sensor index optimization, deep learning classification) are publicly available in the authors' GitHub repository, and the DRONE2REPORT software itself is released as Code · publicThe code and data to reproduce these case studies
can be found at https://github.com/ne1s0n/paper-drone2report (accessed on 13 April
2026).Open asset ↗ne1s0n/paper-drone2reportpdf-page:6 lines:1-59Plant phenotyping relevance match · UnverifiedCrossref · checked 15 Sept 2026
Published17 Apr 2026INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENTCited by 0 · OpenAlex ↗
Abstract- Plant diseases pose a serious threat to global food security by reducing crop yield and quality. To improve detection, a multi-modal diagnostic framework is proposed that combines Vision Transformers (ViTs) for image-based disease classification with Natural Language Processing (NLP) for symptom description and treatment recommendations. The system supports multilingual interaction and generates automatic disease reports, making it accessible to diverse farming communities. By integrating ViT and NLP, the model offers higher diagnostic accuracy and interpretable, farmer-friendly support. Designed for real-world use, it can be deployed on mobile and IoT platforms, enabling smart, interactive decision-making in precision agriculture. Keywords: Plant Disease Detection, Vision Transformers, NLP, Multi-Modal Learning, Precision Agriculture.
Why it matches plant phenotyping methods植物画像から病害を分類するマルチモーダル診断手法が研究の中心であり、植物の病害状態を直接推定するため、フェノタイピング手法として採用する。
abstracta multi-modal diagnostic framework is proposed that combines Vision Transformers (ViTs) for image-based disease classification
Plant phenotyping relevance match · UnverifiedOpenAlex · Crossref · Europe PMC · checked 5 Sept 2026
High-throughput plant phenotyping (HTPP) is increasingly limited by the mismatch between the need for field-relevant, fine-grained phenotypic information and the restricted capability of conventional observation platforms under complex agricultural conditions. Ground mobile robots are emerging as the key carrier for resolving this gap because they combine close-range sensing, autonomous mobility, and physical interaction within real field environments. In this paper, a structured scoping review is presented using a closed-loop perception-decision-action pipeline as the organizing principle. Within this framework, recent advances are synthesized from the perspectives of multimodal fusion, localization-aware sensing, motion planning, deep-learning-based phenotypic analysis, active observation, robotic intervention, and edge deployment. The review further clarifies the complementary roles of Unmanned Aerial Vehicles (UAVs), Unmanned Ground Vehicles (UGVs), and air-ground collaboration in multiscale phenotyping workflows. Beyond summarizing technologies, the article provides three concrete deliverables: a structured taxonomy of mobile phenotyping systems; comparative tables covering sensing modalities, localization/navigation methods, and AI models; and a research agenda linking technical progress to field deployability. The synthesis highlights four persistent bottlenecks, namely environmental generalization, annotation scarcity, limited standardization and reproducibility, and the gap between advanced models and agricultural edge hardware. Overall, ground robots are identified not merely as sensing platforms, but as the central system architecture for advancing mobile phenotyping toward autonomous, fine-grained, and field-deployable operation.
Why it matches plant phenotyping methods植物フェノタイピング用移動ロボットについて、センシング、表現型解析、プラットフォーム分類、比較、標準化を中心に扱う方法論レビューであり、対象範囲に明確に合致する。
abstractIn this paper, a structured scoping review is presented using a closed-loop perception-decision-action pipeline as the organizing principle.
Abstract: Crop diseases significantly reduce agricultural productivity and directly impact farmers’ income. Early detection of plant diseases is essential for effective crop management and the prevention of large-scale crop loss. This paper proposes an IoT-based crop disease detection system using deep learning techniques for automatic identification of plant diseases from leaf images. The system utilizes a Convolutional Neural Network (CNN) model to classify crop diseases from images captured through an IoT camera module [1], [2].To improve interpretability, the Grad-CAM technique is used to generate heatmaps highlighting infected regions on the leaf [4]. Additionally, the system provides automated recommendations for disease management and communicates them to farmers through both GSM-based SMS alerts and a speaker module [11], [16].The speaker module converts prediction results into audio output in the farmer’s native language, announcing the detected disease, recommended pesticide, duration of application, and severity level. This feature enhances accessibility for illiterate and regional-language users. The proposed system integrates image processing, machine learning, IoT devices, and multimodal communication technologies to provide a scalable and practical solution for precision agriculture [9], [10]. Index Terms: Crop Disease Detection, Internet of Things (IoT), Deep Learning, Convolutional Neural Network (CNN), Grad-CAM, Precision Agriculture, GSM Communication, Speaker Module, Audio Output, Smart Farming.
Why it matches plant phenotyping methods葉画像から植物病害と感染領域・重症度を推定する画像解析手法が研究の中心であり、植物状態の表現型推定に該当する。
abstractThis paper proposes an IoT-based crop disease detection system using deep learning techniques for automatic identification of plant diseases from leaf images.
Introduction In protected horticultural production, early disease identification and precise intervention are critical for safeguarding crop yield and quality while reducing chemical pesticide inputs. However, early-stage greenhouse diseases often exhibit extremely subtle visual symptoms, and their occurrence and progression are highly dependent on environmental condition variations, making stable and reliable early warning difficult to achieve using conventional methods based on single visual information or simple multimodal fusion. Methods To address this challenge, a visual-environment joint early disease perception framework for greenhouse horticultural crops is proposed. Through an environment-guided visual attention mechanism and a spatial-temporal joint modeling strategy, environmental variables such as temperature, humidity, vapor pressure deficit, and CO 2 concentration are transformed from passive features into active priors, thereby guiding visual feature learning and enhancing sensitivity to weak disease signals. The proposed method is systematically validated on a real-world greenhouse multimodal temporal dataset. Results Experimental results demonstrate that the proposed approach achieves an accuracy of 91.3%, a recall of 88.9%, and an F1-score of 89.8% in overall disease detection tasks, significantly outperforming multiple baseline models based on convolutional neural networks (CNNs), Transformers, and existing multimodal fusion strategies. In early-stage disease detection scenarios, early precision and early recall reach 88.5% and 86.1%, respectively, with the lead time extended to 2.7 days, indicating a clear advantage in early warning capability. Ablation studies further verify the critical roles of environment-guided attention, spatial-temporal joint modeling, and the joint loss function in improving early detection performance and stability. Discussion This study provides a practically valuable technical pathway for early intelligent warning and precise regulation of greenhouse crop diseases. By integrating environmental dynamics with visual perception, the proposed framework improves the sensitivity and robustness of early disease detection in complex greenhouse conditions, showing strong potential for practical deployment in protected horticulture.
Why it matches plant phenotyping methods環境情報と画像・時系列情報を統合して、植物病徴を早期検出する深層学習フレームワークの開発と実データでの系統的検証が中心であり、植物の病害状態を直接推定するため。
abstracta visual-environment joint early disease perception framework for greenhouse horticultural crops is proposed.
Abstract In sustainable agriculture, detecting pests and diseases early is critical. Recent technological advances in deep learning (DL) and multimodal imaging like multispectral and thermal data crop health monitoring is promising. Despite the progress, obtaining high accuracy across various crops with real-time performance is still a challenge. The hybrid convolutional neural network (CNN)-attention model integrating multispectral and thermal data for pest and disease detection has been introduced. A total of 1760 samples were collected from six crops (maize, rice, wheat, tomato and cassava), across different growth stages, labelled fungal, bacterial, viral and pest infections. The data was divided into 70% training, 15% validation, and 15% test sets. 3,500 samples were used for training. 750 samples were used for validation and test set. The hybrid CNN-attention model was contrasted with certain baseline models (SVM, Random Forest, CNN-RGB, CNN-Multispectral) and certain fusion methods (early, late, and hybrid fusion) based on accuracy, precision, recall, F1-score, and early detection sensitivity. The highest accuracy of 91.0% for rice at the vegetative stage was achieved by the hybrid model. It beats baseline and fusion models. The F1-score of the classification was reasonably high. Rice's sensitivity is 88.1%, and maize is 87.3%. The model fared well for all classes, getting 92.0 % for the healthy plant and 88.2 % for pest infestation. Future work can enhance the dataset with more crops and diseases and environmental factors and optimize detection time and early sensitivity for real-time deployment in agricultural decision support systems.
Why it matches plant phenotyping methodsマルチスペクトル・熱画像から植物の病害および害虫状態を推定するCNNモデルを開発し、複数モデルとの比較検証を行っており、表現型取得・判定手法が中心である。
titleUsing Multispectral Imaging and Artificial Intelligence to Detect Crop Diseases and Pests Early
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 5 Sept 2026
Plant diseases remain a major threat to global food production, causing significant yield losses and economic impact worldwide. Early and precise disease detection is crucial for effective crop management, yet conventional diagnostic approaches are often slow, labor-intensive, and rely on specialized expertise that may not be widely accessible. Recent advances in artificial intelligence (AI), particularly deep learning–based image analysis, offer scalable and automated solutions for plant disease recognition. This review critically examines forty-one peer-reviewed studies published between 2008 and 2025, selected following PRISMA guidelines from major scientific databases. We summarize key methodological developments, including convolutional neural networks, vision transformers, transfer and few-shot learning, and multimodal sensing approaches, highlighting their reported performance and limitations. Although many models achieve high accuracy in controlled datasets, their effectiveness often decreases under real-field conditions due to environmental variability, limited training data, and practical deployment constraints. We discuss existing challenges and propose future research directions, emphasizing improved robustness in field environments, development of lightweight and explainable models suitable for edge deployment, and integration with precision agriculture systems. This review aims to guide the design of reliable, practical, and scalable AI-driven plant disease detection strategies.
Why it matches plant phenotyping methods植物病害の画像認識・マルチモーダルセンシング手法を体系的にレビューしており、病害状態の表現型推定法が中心である。
abstractThis review critically examines forty-one peer-reviewed studies published between 2008 and 2025, selected following PRISMA guidelines from major scientific databases.
Introduction Accurate identification of rice seedling age is essential for guiding precise field management and optimizing agronomic practices. However, traditional identification methods mainly rely on manual experience or simple visual cues and often lack robustness under complex field conditions such as illumination variation, background interference, and subtle morphological differences between adjacent growth stages. Therefore, developing a reliable and automated method for fine-grained recognition of rice seedling stages is of great importance. Methods To address this problem, this study proposes two deep learning models for automatic recognition of 13 rice seedling stages. The first model, Lresnet50, enhances visual feature representation by improving the baseline Resnet50 with a Row-Prior Strip Attention (RPS) mechanism, a Feature Pyramid Network (FPN) for multi-scale feature extraction, and Dynamic Channel Pruning (DCP) to reduce redundant channels and improve computational efficiency. Based on this model, a multimodal framework named M-Lresnet50 is further developed by integrating image features with temporal environmental data through a Long Short-Term Memory (LSTM) network, enabling cross-modal feature fusion and improving recognition of continuous seedling growth stages. Results Experimental results demonstrate that the proposed models achieve high accuracy in recognizing 13 rice seedling stages. The Lresnet50 model achieves an average classification accuracy of 97.70%, outperforming several existing convolutional neural network architectures and showing strong performance in transitional growth stages where morphological differences are subtle. By integrating visual features with temporal environmental information, the multimodal M-Lresnet50 further improves the accuracy to 98.33%. The model contains 27.656 million parameters with a computational complexity of 13.965 GFLOPs, indicating a good balance between recognition accuracy and computational cost. Discussion The results confirm the effectiveness of the proposed improvements and multimodal fusion strategy. The Row-Prior Strip Attention (RPS) enhances the model's ability to focus on row-structured crop regions, while the Feature Pyramid Network (FPN) improves multi-scale feature representation. In addition, Dynamic Channel Pruning (DCP) reduces redundant channels and improves computational efficiency. The integration of temporal environmental information through the multimodal framework further enhances the robustness and consistency of seedling stage recognition. Overall, the proposed approach provides a practical solution for intelligent monitoring of rice seedling growth in greenhouse environments.
Why it matches plant phenotyping methods画像と環境時系列データからイネ幼苗の生育段階を自動認識する深層学習手法を開発し、複数モデルとの精度比較も行っているため、植物状態の取得・推定が中心である。
abstractthis study proposes two deep learning models for automatic recognition of 13 rice seedling stages.
Abstract Plant diseases are a major challenge in agriculture, leading to reduced crop yield and economic losses. Traditional methods of disease detection rely on manual inspection, which is time-consuming and less accurate. With the advancement of Deep Learning, Computer Vision, and Internet of Things, automated plant disease detection systems have been developed to improve accuracy and efficiency. Recent research shows that deep learning models, especially Convolutional Neural Networks (CNNs), are widely used for plant disease classification due to their ability to automatically extract features from images and achieve high accuracy [3], [34] . Transfer learning techniques using models such as ResNet and VGG further enhance performance, particularly when datasets are limited [9], [41] . In addition, lightweight models like MobileNet and ShuffleNet are being developed for deployment on mobile and IoT devices [2], [12] . Many studies use benchmark datasets such as PlantVillage, which provide high accuracy under controlled conditions. However, real-world applications face challenges such as varying lighting conditions, complex backgrounds, and limited dataset diversity [13], [15] . To address these issues, researchers are exploring advanced techniques such as data augmentation, object detection models like YOLO, and multimodal approaches that combine image data with environmental sensor data [5], [10] . IoT-based systems are also gaining importance as they enable real-time monitoring of crop conditions using sensors and smart devices [21], [26] . Furthermore, Explainable AI techniques such as Grad-CAM are being used to improve model transparency and help users understand the prediction results [25] . Despite significant progress, challenges remain in terms of model generalization, computational complexity, and real-time deployment. Future research should focus on developing lightweight, efficient, and scalable models, along with the use of large real-field datasets and edge computing technologies. Keywords Plant Disease Detection; Leaf Image Analysis; Convolutional Neural Network (CNN); Deep Learning; Transfer Learning; Computer Vision Internet of Things; Lightweight Models; Object Detection (YOLO); Multimodal Data Fusion; Explainable Artificial Intelligence (XAI)
Why it matches plant phenotyping methods植物病害を葉画像から推定するCNN・画像解析手法を体系的に扱うレビューであり、植物の病態を観測・推定するフェノタイピング手法が中心です。
titleHybrid CNN-Based Plant Disease Detection for Smart Agriculture: A Review
Field / plotMultimodalLeafSegmentationDisease symptoms / severity
Camellia oleifera is an economically vital woody oil crop. Its productivity and oil quality are severely compromised by various diseases. Implementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection. Despite recent progress, existing segmentation methods struggle with three primary challenges: semantic ambiguity arising from evolving pathological stages, blurred boundaries due to overlapping lesions, and the high omission rate of micro-lesions. To address these issues, this paper presents TB-DLossNet (Text-Conditioned Boundary-Aware Network with Dynamic Loss Reweighting), a novel segmentation framework based on semantic-visual multi-modal fusion. Leveraging VMamba as the visual backbone, the proposed model innovatively integrates BERT-encoded structured text as an auxiliary modality to resolve visual ambiguities through cross-modal semantic guidance. Furthermore, a boundary enhancement branch is incorporated alongside a multi-scale deep supervision strategy to mitigate boundary displacement and ensure the topological continuity of lesion structures. To tackle the detection of small-scale targets, we designed a dynamic weight loss function conditioned on lesion area, significantly bolstering the model's sensitivity to minute pathological features. Additionally, to alleviate the scarcity of high-quality data, we curated a comprehensive multi-modal dataset encompassing seven typical diseases of Camellia oleifera . Experimental results demonstrate that TB-DLossNet achieves a Mean Intersection over Union (mIoU) of 87.02%, outperforming the state-of-the-art unimodal VMamba and multimodal Lvit by 4.9% and 2.59%, respectively. Qualitative evaluations confirm that our model exhibits lower false-negative rates and superior boundary-fitting precision in heterogeneous field scenarios. Finally, generalization tests on an apple disease dataset further validate the robustness and transferability of the proposed framework.
Why it matches plant phenotyping methods植物病害の病斑を画素レベルで抽出する新規セグメンテーション手法を開発し、データセット整備と性能比較・汎化検証も行っているため、病害状態の画像ベース表現型計測が中心である。
abstractImplementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection.
Reproduction assets foundThe authors state their code and experimental dataset (the multimodal Camellia oleifera disease segmentation dataset) are publicly available on GitHub, matching an allowed URL.Code · publicOur code and experimental dataset are available at https://github.com/zzzsq239/TB-1.Open asset ↗zzzsq239/TB-1html-lines:820-841Plant phenotyping relevance match · UnverifiedCrossref · checked 15 Sept 2026
Global food security faces unprecedented pressure from crop diseases, which are responsible for annual yield losses estimated at 20–40% of all food production. The application of deep learning methodologies, particularly convolutional neural network (CNN) architectures, to the automated detection and classification of crop diseases has emerged as a transformative paradigm within the domain of smart agriculture. A structured literature search was conducted using the academic databases Web of Science, Scopus, Google Scholar, and PubMed, covering the publication period from January 1996 to March 2026. The search strategy employed a combination of controlled vocabulary and free text search strings, including but not limited to the following terms and their Boolean combinations. The review examining the theoretical underpinnings and empirical performance of diverse CNN architectures including VGGNet, ResNet, Inception, DenseNet, EfficientNet, and Vision Transformers as applied to plant pathology. Special attention is directed towards the role of multispectral and hyperspectral imaging modalities, which extend disease detection capabilities beyond the visible spectrum and enable the identification of latent biochemical stress signatures before visible symptom onset. The review further explores the critical contribution of transfer learning in addressing the perennial challenge of limited annotated agricultural datasets, demonstrating how pre trained models can be fine tuned to achieve high diagnostic accuracy across diverse crop pathogen combinations. A dedicated section examines the rapidly maturing field of Explainable AI (XAI), with particular focus on gradient weighted class activation mapping (Grad CAM), integrated gradients, and SHAP based methods, which are essential for building agronomist trust and regulatory acceptability. The synthesis identifies persistent challenges including domain shift, class imbalance, computational constraints in field deployable systems, and the scarcity of standardised benchmark datasets. The review concludes with a forward looking perspective on federated learning, multimodal fusion architectures, and the integration of UAV based sensing with edge computing as the frontier of next generation agricultural AI systems.
Why it matches plant phenotyping methods植物病害の画像ベース検出・分類手法を中心に、CNN、マルチスペクトル/ハイパースペクトル画像、説明可能AIなどをレビューしており、植物状態(病害)の取得・推定方法が主題である。
abstractThe review examining the theoretical underpinnings and empirical performance of diverse CNN architectures including VGGNet, ResNet, Inception, DenseNet, EfficientNet, and Vision Transformers as applied to plant pathology.
Plant diseases pose a significant threat to global agriculture, impacting crop yields and quality. Early and accurate detection is essential for effective health management but remains challenging due to visual similarity among diseases and complex field backgrounds. This study introduces AgriMM, a novel multi-modal detection framework that integrates visual images with expert-validated textual descriptions to improve diagnostic precision. The framework features three key innovations: a Hybrid Convolutional-Attention Collaborative Backbone (HCACB) to capture both fine-grained lesions and global context; a Context-enhanced Visual-Language Path Aggregation Network (CVL-PAN) for multi-scale feature fusion; and an Adaptive Region-Text Contrastive Learning (AR-TCL) module to enforce precise semantic alignment. We constructed a comprehensive dataset comprising 30,000 images and detailed symptom descriptions across five major crops (tomato, cucumber, pepper, eggplant, and squash). Experimental results demonstrate that AgriMM achieves a mean Average Precision (mAP) of 95.2%, significantly outperforming state-of-the-art unimodal baselines by 11.6%. These findings confirm that integrating linguistic semantic priors effectively resolves visual ambiguity, providing a robust tool for precision agriculture and sustainable crop protection.
Why it matches plant phenotyping methods植物病害の症状を画像から検出・診断するマルチモーダル手法を開発し、データセットと性能比較で検証しているため、植物フェノタイピング手法が中心である。
abstractThis study introduces AgriMM, a novel multi-modal detection framework that integrates visual images with expert-validated textual descriptions to improve diagnostic precision.
Existing tomato datasets often focus on short-term experiments or lack integrated environmental and agronomic data. We present Horti-M3-Tomato, a comprehensive three-year dataset collected in Northeast China's greenhouse, including high-resolution RGB images, environmental sensor data (recorded every 30 minutes), soil conditions, and detailed agronomic records such as yield data and management practices. Spanning three growing seasons (2023-2025), the dataset integrates temporal imaging, environmental monitoring, soil data, and manual phenotypic and yield records. Horti-M3-Tomato supports research on growth dynamics, genotype-environment interactions, and provides a benchmark for AI-based phenotyping and precision horticulture. The dataset is openly available for further research in controlled-environment agriculture.
Why it matches plant phenotyping methodsトマトの画像・環境センサーデータ・手動表現型記録を統合したデータセットであり、AIベースの表現型解析のベンチマークとして明示されているため、表現型データ基盤が中心です。
abstractincluding high-resolution RGB images, environmental sensor data (recorded every 30 minutes), soil conditions, and detailed agronomic records such as yield data and management practices.
Accurate and non-destructive evaluation of grape quality is crucial for intelligent viticulture, yet most existing approaches address cultivar classification and soluble solid content (SSC) prediction as independent tasks based on single-modality data, limiting robustness and practical applicability. This study proposes DualStream-RTNet, a unified multimodal deep learning framework that simultaneously performs grape cultivar classification and SSC prediction by integrating RGB-HSV fused images and PCA-compressed hyperspectral spectra. The dual-stream architecture enables the complementary learning of external chromatic-textural cues and internal physicochemical information, while a Transformer-enhanced fusion module strengthens global representation and cross-modal correlation. A dataset of 864 berries from five grape cultivars was used to validate the model. DualStream-RTNet achieved 93.64% classification accuracy, outperforming ResNet18 and other CNN baselines, and produced more compact and consistent confusion-matrix patterns. For SSC prediction, it consistently yielded the highest performance across cultivars, with R2p values up to 0.9693 and RMSE as low as 0.2567, surpassing the PLSR, SVR, LSTM, and Transformer regression models. These results demonstrate the superiority of the proposed framework in capturing both visual and spectral characteristics. DualStream-RTNet provides an efficient and scalable solution for comprehensive grape quality assessment, offering strong potential for real-time sorting, precision grading, and smart agricultural applications.
Why it matches plant phenotyping methodsRGB-HSV画像とハイパースペクトルを統合し、ブドウ果実のSSCを非破壊推定する新規モデルを開発・検証しており、果実形質の取得・推定法が研究の中心である。
abstractThis study proposes DualStream-RTNet, a unified multimodal deep learning framework that simultaneously performs grape cultivar classification and SSC prediction by integrating RGB-HSV fused images and PCA-compressed hyperspectral spectra.
Traditional methods for identifying salt tolerance levels in soybean varieties are often cumbersome, time-consuming, and labor-intensive. These challenges are further exacerbated by the limited utility of chlorophyll fluorescence imaging phenotype data, which are insufficiently diverse and difficult to analyze. Additionally, the corresponding parameter text data have not been fully explored and utilized. In this study, salt stress experiments were conducted on 178 soybean varieties, and a multimodal dataset comprising chlorophyll fluorescence images and corresponding textual data was constructed using a chlorophyll fluorescence imaging instrument. A novel gated mechanism network for learnable image-text interaction (Mm-VitnNet) is proposed, which enables global cross-modal interaction between image and text data. The model introduces a gated mechanism to dynamically regulate the fusion intensity of cross-modal information and incorporates two learnable tokens that focus on feature learning for each individual modality. This approach effectively mitigates interference between modalities while preserving modality-specific features, thereby enhancing model performance. The proposed model demonstrates an accuracy rate of 98.97%, significantly outperforming typical models: it improves by 1.09 and 2.33 percentage points compared to CNN-based models such as EfficientNetV2-s (97.88%) and MobileNetV2 (96.64%), respectively, and by 3.21 and 2.60 percentage points compared to Transformer-based Swin Transformer_tiny (95.76%) and hybrid models like MobileViT_S (96.37%), respectively. The model has 10.22M parameters and a computational cost (FLOPs) of 1.84G, which is significantly lower than models like VGG and ResNet50, and only slightly higher than some lightweight CNNs, achieving an effective balance between accuracy and efficiency. The improved model demonstrates notable performance in identifying samples with varying salt tolerance levels, even under limited computational resources, ensuring reliable classification performance. Moreover, this multimodal non-destructive identification method based on chlorophyll fluorescence technology offers an efficient and feasible approach for assessing the salt tolerance levels of soybeans, while also advancing agricultural phenotyping towards greater precision and intelligence.
Why it matches plant phenotyping methodsダイズの塩耐性という植物状態をクロロフィル蛍光画像から推定するマルチモーダル画像解析手法を開発・評価しており、表現型取得と分類モデルが研究の中心である。
abstracta multimodal dataset comprising chlorophyll fluorescence images and corresponding textual data was constructed using a chlorophyll fluorescence imaging instrument.
Plant phenotyping relevance match · UnverifiedOpenAlex · Europe PMC · Crossref · checked 15 Sept 2026
High-throughput plant phenotyping is increasingly constrained by the mismatch between the demand for field-relevant, fine-grained phenotypic data and the limited capability of conventional observation platforms under complex agricultural conditions. In this context, mobile phenotyping systems, particularly ground robots, are emerging as a key technological pathway for bridging macro-scale monitoring and organ-level trait analysis. This review examines the development of mobile phenotyping platforms for high-throughput plant phenotyping, with emphasis on the evolving role of ground robots in field-based sensing, decision-making, and active interaction. We first compare the functional characteristics of unmanned aerial vehicles and unmanned ground vehicles and discuss their complementarity in multiscale phenotypic data acquisition. We then summarize recent advances in the core technical framework of mobile phenotyping robots, including multimodal perception, localization and mapping, motion planning, deep-learning-based phenotypic analysis, active observation, robotic intervention, and edge deployment. Major challenges are further discussed, particularly those related to environmental generalization, data annotation, standardization, reproducibility, and long-term field reliability. Finally, future directions are outlined from the perspectives of air–ground collaboration, multi-robot systems, foundation models, and embodied intelligence. This review highlights ground robots as a central carrier for advancing mobile phenotyping toward autonomous, fine-grained, and field-deployable systems.
Why it matches plant phenotyping methods植物フェノタイピング用の移動ロボットプラットフォームと、マルチモーダルセンシング・表現型解析などの技術を中心に扱うレビューであり、対象分野の方法論的レビューに該当する。
abstractThis review examines the development of mobile phenotyping platforms for high-throughput plant phenotyping
Pest and pathogen pressure in potato cultivation is increasingly affecting the potato quality and yield. The Netherlands, as the largest seed potato producer around the world, is particularly threatened by blackleg disease and potato virus Y (PVY). Uncrewed aerial vehicle (UAV)-based imaging combined with machine- and deep-learning methods have shown clear potential for potato disease identification, offering advantages over conventional human inspections, which are labor-intensive, expertise-demanding, and often subjective. Most existing studies focused on RGB data and pixel-level classification, producing maps that have limited practical value for targeted removal of infected plants. Earlier work demonstrated the potential of plant-level disease detection approaches. For example, Jia[1] employed hyperspectral data (specifically the first three principal component analysis (PCA) bands) with a YOLOv5s model to distinguish the blackleg- and PVY-infected plants from healthy ones, yielding average mAP@.50 scores of 0.85 for blackleg detection and 0.82 for PVY detection. Gibson-Poole[2] applied object-based image analysis (OBIA) to detect blackleg disease with RGB imagery, achieving a total accuracy of 87%. The findings suggest that multi-modal data (combining hyperspectral and RGB imagery) hold strong potential for plant-level disease detection. We aim to identify the most informative features derived from hyperspectral data and to investigate their integration with RGB data to enhance potato disease detection performance.We proposed early fusion (E), where data were concatenated channel-wise before network input, and middle fusion (M) architectures, where features were extracted separately within a two-branch network and then merged at an intermediate stage, to integrate hyperspectral features and RGB imagery for potato disease detection. To reduce hyperspectral dimensionality, two feature sets were extracted: (i) the first three PCA bands, and (ii) 10 vegetation indices (VIs) selected from 64 candidates using variance inflation factor analysis to mitigate multicollinearity. Consequently, four models were developed and evaluated: E-PCA-RGB, E-VI-RGB, M-PCA-RGB, and M-VI-RGB. Unlike previous studies that focused on a single disease, our models detected blackleg-infected, PVY-infected, and healthy plants simultaneously. E-VI-RGB achieved the highest mAP@.50 value of 86.65±1.53, followed by M-VI-RGB (85.74±1.75). E-PCA-RGB and M-PCA-RGB yielded mAP@.50 scores of 83.00±2.52 and 83.11±2.20, respectively. These results demonstrate that combining hyperspectral features with RGB imagery improves detection performance compared with single-modality approaches (RGB 83.21±1.31, PCA 79.71±1.30, VIs 85.31±2.11). Our findings highlight the potential of multimodal fusion for potato disease detection in practice. The methods could enable automated systems not only to identify infected plants but also to support timely removal with machinery, mitigating the spread of disease in potato fields. The generalizability of our approach will be further tested and analyzed in future work.References[1] Jia, T., Smigaj, M., Kootstra, G. and Kooistra, L., 2024. Detection of Diseased Potato Plants with UAV Hyperspectral Imagery. In 2024 14th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS) (pp. 1-5). IEEE.[2] Gibson-Poole, S., Humphris, S., Toth, I. and Hamilton, A., 2017. Identification of the onset of disease within a potato crop using a UAV equipped with un-modified and modified commercial off-the-shelf digital cameras. Advances in Animal Biosciences, 8(2), pp.812-816.
Why it matches plant phenotyping methodsUAVのハイパースペクトル・RGB画像からジャガイモ個体の病害状態を推定するマルチモーダル融合手法を開発・評価しており、植物表現型の取得・抽出が研究の中心です。
abstractWe proposed early fusion (E), where data were concatenated channel-wise before network input, and middle fusion (M) architectures, where features were extracted separately within a two-branch network and then merged at an intermediate stage, to integrate hyperspectral features and RGB imagery for potato disease detection.
Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationGrowth / development / phenology
Abstract Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.
Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画のデータセットで、植物フェノロジー自動検出の訓練、検証、ベンチマークを目的とする方法論的成果である。
titleA Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis
Reproduction assets foundThe paper is a Data Note describing the Multi-Modal Actinidia chinensis Phenology Dataset, which is explicitly stated to be publicly available on Zenodo with a DOI matching an allowed URL. The dataset contains the paper's own phenotyping assets: 1,665 annotated images with bounding-box phenological labels, georeferenedDataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available
at Zenodo: https://doi.org/10.5281/zenodo.17371025. This dataset comprises two components: (1) 1 665
JPEG images (1 024 × 1 024 pixels) with corresponding Pascal VOC XML annotation files containing bounding
box coordinates and phenological class labels, and (2) 24 MP4 video files (3 840 × 2 160 pixels) with corresponding
GPX coordinate files and Excel validation files containing manual ground truth counts.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:13 lines:1-62Plant phenotyping relevance match · UnverifiedCrossref · checked 15 Sept 2026
MultimodalFruitLeafClassificationObject detectionPhysiological trait estimationGrowth / development / phenologyFruit / seed / panicle traitsPlant / canopy temperature
ABSTRACT Plant‐wearable sensors are essential for real‐time, in situ monitoring in smart agriculture, yet their adoption is constrained by functional simplicity and inadequate mechanical compliance. Herein, we present a conformable plant‐attachable Janus electronic skin (SPM/HMA@P) for multimodal phenotyping. Featuring a scenario‐specific encapsulation strategy fabricated via a bottom‐up approach, the device enables non‐invasive monitoring of multiple physiological parameters. The sensor integrates a conductive base composed of biocompatible quaternized chitosan, multi‐walled carbon nanotubes, and silver nanowires, a gas‐sensing functional layer, and an adhesive polydimethylsiloxane encapsulation. In the fully encapsulated configuration, SPM/HMA@P exhibits a 210.9% fracture elongation and an adhesion force exceeding 0.3 N on leaves, alongside excellent biocompatibility. This configuration allows for simultaneous monitoring of leaf temperature and growth‐induced strain, maintaining stable signals over 5 days. Conversely, the semi‐open configuration demonstrates high ethylene sensitivity (0.5–100 ppm, theoretical detection limit of 0.12 ppm), with response and recovery times of 300 and 120 s, and a lifespan of over 10 days. Coupled with a convolutional neural network (CNN), it achieves 96.8% accuracy in classifying ethylene signals from fruits. This robust multimodal platform addresses key challenges in plant physiological monitoring, holding great potential for smart crop breeding, postharvest assessment, and phenotyping.
Why it matches plant phenotyping methods植物に装着するマルチモーダルセンサーを開発し、葉温度・成長誘導ひずみ・エチレンなどの生理状態を非侵襲的に測定するプラットフォームであり、植物フェノタイピング手法が中心的に扱われている。
titleConformable Plant‐Attachable Janus E‐Skin with Sandwich Architecture for Plant Multimodal Phenotyping
Agriculture currently faces the dual pressures of ensuring global food security and adapting to rapid climate change. To cope with these challenges, researchers have introduced several modern mechanization technologies, including advanced farm machinery, autonomous navigation systems, artificial intelligence, sensing technologies, and communication tools, to enhance productivity and sustainability (Syed et al., 2025a). These technologies enable data-driven decision-making by allowing continuous, large-scale acquisition and analysis of crop and environmental information. Consequently, accurately predicting crop yields and monitoring plant health in real time have become critical prerequisites for precision agricultural management (Syed et al., 2025). Traditional measurement methods-often labor-intensive, destructive, and spatially limited-are increasingly unable to meet the demands of modern large-scale farming. In this context, the integration of Remote Together, these ten contributions illustrate the maturation of agricultural remote sensing, moving towards models that are not only more accurate but also lighter, more interpretable, and more resilient to environmental noise. By combining satellite and UAV data with advanced computational models, these innovative approaches are paving the way for a more resilient and productive global food system. Future research will increasingly focus on improving the precision of crop yield estimation models through multi-dimensional analyses. As agricultural environments grow more complex, integrating AI-powered models with multi-sensor fusion technologies will be essential. Innovations such as lightweight neural networks and multimodal cross-attention frameworks will enable the detection of small, occluded, and densely packed targets with greater accuracy, thereby refining crop-specific metrics such as photosynthetically active radiation (FPAR) and nitrogen content. This, in turn, will enhance crop health monitoring and yield predictions.Additionally, UAV-based remote sensing, combined with multitier feature selection, will improve nitrogen content analysis in crops such as cotton, while image dehazing models and light-use efficiency frameworks will bolster biomass estimation.Emerging technologies such as the Ta-YOLO framework will further optimize small fruit detection in dense canopies, advancing overall crop detection accuracy.A key challenge lies in adapting these models to handle real-world complexities, such as variable environmental conditions. Future work will focus on improving the robustness of these models through dynamic coding networks and performance optimization, ensuring they can operate in heterogeneous agricultural environments.Interdisciplinary collaboration between agriculture, AI, and remote sensing experts will accelerate the development and deployment of these approaches, paving the way for more efficient crop yield estimation systems that are critical for ensuring food security and sustainable agricultural practices.
Why it matches plant phenotyping methods作物収量・健康・バイオマス・窒素含量などの植物形質を、衛星・UAVリモートセンシングと計算モデルで推定する手法群を中心に扱う編集レビューであり、方法論的役割が明確。
titleInnovative approaches in remote sensing for precise crop yield estimation: advancements, applications, and future directions
Controlled-environment agriculture (CEA) and circular production systems require coordinated monitoring of biological and physicochemical processes across trophic levels. This project report presents the implementation of a multi-trophic controlled-environment agriculture demonstrator that integrates computer-vision-based monitoring with established sensor infrastructure for aquaculture, poultry, plants, microalgae, duckweed, and insect modules. Stereo imaging and RGB-D systems are deployed for non-invasive quantification of fish biomass and plant growth, while continuous water-quality and environmental measurements (e.g., pH, dissolved oxygen, nitrate, ammonium, temperature, CO2) provide complementary process data. These data streams are synchronized within a shared database architecture to enable cross-module evaluation of nutrient dynamics, growth progression, and operational stability under real facility conditions. The implemented framework demonstrates how computer vision can extend conventional sensor-based monitoring by directly capturing biological performance indicators across aquatic, terrestrial, and microbial domains. While advanced predictive modeling and full digital twin simulation remain future development steps, the realized data-integration architecture establishes a structural foundation for the systematic evaluation of circular indoor food-production systems. The demonstrator illustrates how multimodal monitoring can support nutrient recirculation, transparency of biological variability, and data-driven assessment within controlled multi-trophic environments.
Why it matches plant phenotyping methodsステレオ画像およびRGB-Dによる植物生長の非破壊定量と、データ統合型モニタリング基盤の実装が中心的に記述されており、植物フェノタイピング基盤として収載対象です。
abstractStereo imaging and RGB-D systems are deployed for non-invasive quantification of fish biomass and plant growth
Abstract Plant diseases are a serious issue that cause food shortages and financial losses. In large-scale farming, traditional disease detection techniques that primarily rely on expert inspection are frequently unfeasible. Deep learning and super-resolution methods for plant disease detection are thoroughly reviewed in this study, with an emphasis on their use in leaf-level imaging and UAV-based monitoring. This thorough review was carried out using the PRISMA framework, looking at peer-reviewed publications from important databases like Google Scholar, ScienceDirect, Web of Science, and Scopus. According to the review, low spatial resolution, environmental variability, occlusion, and domain shift cause convolutional and transformer-based models to perform poorly at canopy and field scales, despite achieving high accuracy on leaf-level datasets. Despite improving the perceptual quality of aerial imagery, super-resolution techniques are still difficult to incorporate into disease detection pipelines because of their high computational overhead, lack of task-aware training, scarcity of annotated UAV datasets, and poor generalization in real-world scenarios. There are few actual architectural implementations of cross-scale integration strategies currently in use; most of them are conceptual. Deep learning for plant disease detection has come a long way, but reliably deploying this technology outside of controlled environments remains challenging due to scale differences and data availability constraints. Coordinated developments in cross-scale learning approaches, data collection, and model design are needed to address these issues. The results also emphasize the significance of multimodal data fusion, super-resolution-assisted domain adaptation, and hierarchical transfer learning as viable approaches to enhancing the scalability and dependability of plant health monitoring systems. This review highlights useful research directions and provides a critical overview of current methods. It encourages the development of AI-powered solutions for better crop management, early disease detection, and sustainable farming methods.
Why it matches plant phenotyping methods植物病害の葉・キャノピー画像から病害状態を検出する深層学習および超解像手法を対象とした体系的レビューであり、植物状態の取得・推定手法が中心です。
titleA systematic review of deep learning and super resolution techniques for leaf level and canopy level plant disease detection
Plant pathogens cause yield losses worldwide, threatening food security and livelihoods. Because early infection is difficult to diagnose, management often relies on prophylactic pesticide use, increasing costs and environmental impact. Here we present PSNet, a multimodal framework that fuses hyperspectral imaging with RGB information for presymptomatic plant disease detection, together with a low-cost hyperspectral camera incorporating a 3D-printed housing, costing under £500. We validate PSNet using Arabidopsis thaliana infected with the oomycete Albugo candida . Imaging at 2 and 4 days post inoculation, prior to visible symptoms, revealed spectral signatures that distinguished infected from healthy plants, while imaging at 6 days post inoculation captured the transition toward early symptom emergence. Discriminative spectral regions overlapped wavelengths associated with plant responses to biotic stress, supporting the biological plausibility of these signatures. Performance was evaluated using strict plant-level partitioning, ensuring samples from the same plant were confined to a single split. On a four-class task (healthy, 2 dpi, 4 dpi, 6 dpi), PSNet achieved 90.00% accuracy and 97.50% accuracy for binary classification. Together, these results demonstrate that presymptomatic detection is feasible under controlled conditions using low-cost hardware and multimodal learning, underscoring the potential of scalable multimodal systems for early disease monitoring.
Why it matches plant phenotyping methods低コストのハイパースペクトル・RGB融合による植物病害状態の非破壊推定手法を開発し、植物単位で性能検証しているため、フェノタイピング手法が中心である。
abstractHere we present PSNet, a multimodal framework that fuses hyperspectral imaging with RGB information for presymptomatic plant disease detection, together with a low-cost hyperspectral camera incorporating a 3D-printed housing, costing under £500.
Unmanned aerial vehicle (UAV) remote sensing has evolved from experimental imaging into an operational diagnostic infrastructure supporting climate-smart agriculture through high-resolution, flexible, and timely crop observation. This review synthesizes advances in UAV platforms, multisensor payloads, artificial intelligence (AI) analytics, and multisource data fusion to evaluate their combined potential for monitoring heterogeneous smallholder systems. A PRISMA-guided analysis of 59 studies (2013–2024) classified sensing architectures, analytical approaches, and application domains across diverse agroecological contexts. Integrated UAV–AI frameworks improve detection of crop stress, yield variability, biomass distribution, and phenological dynamics compared with conventional monitoring, particularly when multimodal sensor data are fused with satellite and ground observations. Predictive performance and diagnostic reliability increase when spectral, thermal, and structural datasets are analyzed jointly using machine-learning or deep-learning models. However, scalability remains constrained by operational, infra-structural, and regulatory factors, especially in resource-limited systems. These findings demonstrate that integrated sensing–analytics systems form a critical foundation for scalable climate-smart agricultural transformation and data-driven decision support across farm, landscape, and institutional scales.
Why it matches plant phenotyping methodsUAVセンシングとAIによる作物ストレス、収量変動、バイオマス、フェノロジーの観測・推定技術を体系的にレビューしており、植物形質・状態の取得方法が中心である。
abstractThis review synthesizes advances in UAV platforms, multisensor payloads, artificial intelligence (AI) analytics, and multisource data fusion
The wheat canopy genome harbors abundant yet untapped genetic variation that could be harnessed to enhance yield potential. The green area index (GAI) is a structural metric that reflects the photosynthetically active canopy surface and is closely linked to final grain yield. Current image-based GAI retrieval methods often suffer from signal saturation and coarse structural depiction, constraining downstream genetic analyses. To address this limitation, we constructed a comprehensive image dataset spanning eight field experiments across China and France, encompassing approximately 600 genotypes under six distinct management regimes. Leveraging this diverse data, we developed a multimodal deep-learning framework augmented by simulated-to-realistic (sim2real) synthetic data transfer. This framework fuses nadir and oblique RGB images with accumulated thermal time to produce high-precision, time-series GAI estimates. Validated on independent testing datasets from both China and France, the multimodal approach demonstrated robust performance with an accuracy of R 2 = 0.88 and an RMSE of 0.49 m 2 m -2 , representing an improvement of about 22% over the traditional gap fraction method. In three site-year field experiments involving 565 genotypes, the GAI dynamics derived from the multimodal approach showed higher broad-sense heritability (0.20-0.48) than those from the gap fraction approach (0.02-0.13) and stronger genotypic correlations with yield (0.19-0.40 versus 0.09-0.31). Furthermore, genetic analysis confirmed the biological fidelity of the estimated traits, identifying loci that co-localize with known architectural regulators such as Rht-D1 , TaTB1-4D , and TaBGC1-4D . Consistently, the multimodal-derived phenotypes were specifically enriched in cell-wall remodeling and hormonal signaling pathways (e.g., brassinosteroid) that directly regulate canopy expansion. Overall, the proposed method offers a powerful tool for unlocking genetic gain in canopy architecture and accelerating canopy-targeted wheat improvement.
Why it matches plant phenotyping methodsデュアルアングルRGB画像と熱時間を統合してGAIを推定する深層学習法を開発し、独立データで検証しているため、植物形質取得法が研究の中心です。
abstractwe constructed a comprehensive image dataset spanning eight field experiments across China and France
Reproduction assets foundThe paper publicly releases its pre-trained multimodal GAI-estimation model weights and inference code on Hugging Face, directly reproducing this paper's phenotyping analysis. The raw image and phenology datasets are not public and require contacting the authors.Code · publicThe pre-trained model weights, inference code, and usage instructions are publicly available in the Hugging Face repository at https://huggingface.co/PheniX-Lab/GAI-Estimation/tree/main .Open asset ↗PheniX-Lab/GAI-Estimationlines:259-277Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
This article describes a multi-sensor dataset collected during the TIRAMISU (Thermal InfraRed Anisotropy Measurements in India and Southern eUrope) campaign at the Nawagam research site in Gujarat, India, during the 2023 monsoon season. The objective was to acquire continuous ground-based optical and thermal measurements over a homogeneous rice canopy across different crop growth stages. The dataset integrates several complementary components. Thermal data were acquired with an Optris longwave infrared camera (8-14 µm) at high temporal resolution, capturing canopy temperature dynamics throughout the diurnal cycle. Optical data were obtained with a Micasense RedEdge-M multispectral sensor, providing imagery in Blue, Green, Red, RedEdge, and Near-Infrared bands with radiometric corrections. An Apogee radiometer supplied reference radiometric temperature. Meteorological measurements included air temperature, humidity, wind speed and direction, and net radiation. Ancillary field measurements comprised Leaf Area Index (LAI), plant height, emissivity sampling, hyperspectral observations, and crop stage information. The datasets are provided with metadata and processing workflows, including calibration procedures for optical reflectance and thermal radiance. Together, these components form a comprehensive record of canopy-atmosphere interactions over a homogeneous rice field. The datasets can support research on optical and thermal directional anisotropy, canopy radiative transfer, emissivity characterization, and crop biophysical parameter estimation. In addition, they are relevant for applications in vegetation monitoring, agricultural water stress assessment, and surface energy balance studies. By combining optical, thermal, and meteorological observations, the resource is suited for multidisciplinary investigations in remote sensing, agronomy, and environmental sciences.
Why it matches plant phenotyping methods光学・熱画像、校正手順、処理ワークフロー、LAIや草丈などの植物形質を含む再利用可能な作物キャノピーデータセットが研究の中心であり、植物表現型取得基盤として適格。
abstractThe dataset integrates several complementary components.
Reproduction assets foundThe paper is a Data in Brief describing the TIRAMISU rice-canopy dataset (thermal/multispectral images, meteorological, ancillary LAI/height, hyperspectral, emissivity) publicly deposited at doi.org/10.6096/1028, including processing scripts (Thermal_CSV_to_Image.py, MicaSense notebook) for reproducibility.Dataset · publicRepository name: Optical, Thermal Infrared, and Meteorological Dataset from the Thermal InfraRed Anisotropy Measurements in India and Southern eUrope (TIRAMISU) Rice Canopy Experiment
Data identification number: doi.org/10.6096/1028
Direct URL to data: https://doi.org/10.6096/1028
Instructions for access: Publicly accessible repository; representative subsets provided with metadata and processing scripts.
Related research article
Pinnepalli, C., Roujean, J.-L., Irvine, M., et al. [ 1 ]. Measuring and modelling directional effects in the frame of TIRAMISU. ISPRS Annals, X–3–2024 , 325–330. https://doi.orgOpen asset ↗doi.org · 10.6096/1028lines:49-77Plant phenotyping relevance match · UnverifiedCrossref · OpenAlex · checked 15 Sept 2026
MaizeRiceSoybeanField / plotMultimodalLiDAR / point cloudRGB / grayscaleMultispectral / hyperspectralThermalWhole plant / canopy / plot / field
Plant phenotyping is essential for elucidating genotype–environment interactions, yet conventional methods remain labor-intensive and low-throughput. TraitDiscover transcends these constraints by uniting multimodal sensing with tightly coupled hardware-software orchestration in a single, end-to-end phenotyping platform. Aligned with the ”Plant Phenotyping Trinity” framework, the system comprises a millimetre-accurate triaxial automation unit, a modular sensor array–RGB imaging, three-dimension laser scanner or LiDAR (3D), infrad (IR) thermal imaging, hyperspectral imaging (HSI), and photosynthesis (PS) imaging–and the dedicated software TraitNavigator suite into one cohesive system. A unified spatiotemporal synchronization mechanism enables robust time-series analysis and fusion of multisource phenotypic data across the entire crop growth period, while the DepthCropSeg algorithm and a night-time imaging module enhance trait extraction under complex conditions, providing G × E × P-ready, multimodal phenotypic datasets. Validation across soybean, maize, and rice trials demonstrated high sensitivity—detecting drought stress four days before visible symptoms, identifying glyphosate injury 24 hours ahead of manual scoring, and quantifying local adaption patterns across ecological gradients. While challenges remain in scaling to complex open-field conditions, TraitDiscover offers a scalable, data-driven approach to accelerate stress phenotyping and breeding decisions and is readily poised for deeper integration with AI to advance sustainable agriculture.
Why it matches plant phenotyping methodsマルチモーダルセンシング、画像解析、同期機構、形質抽出アルゴリズムを統合した植物フェノタイピング基盤の開発と検証が中心であり、ストレス検出や形質定量も実証している。
abstractTraitDiscover transcends these constraints by uniting multimodal sensing with tightly coupled hardware-software orchestration in a single, end-to-end phenotyping platform.
Accurate acquisition of phenotypic characteristics in protected crops is a crucial prerequisite for intelligent control and digital breeding in greenhouses. To accurately assess the phenotypic traits of protected lettuce, a specialized in situ phenotypic detection method has been developed. The Multimodal Features and Attention Mechanism for Phenotype Detection Model (MFAMNet) was developed for protected lettuce, employing a segmented multi-source image dataset for synchronous regression testing. The results revealed that the predicted values generated by MFAMNet exhibited a strong correlation with the measured values, achieving coefficients of determination of 0.96, 0.92, 0.95, 0.94, and 0.95 for plant height, crown width, leaf area, fresh weight, and dry weight, respectively. Ablation tests demonstrated that the deep learning detection framework based on multi-modal feature fusion significantly outperformed single-feature detection models, highlighting the advantages of integrating diverse data modalities. In addition, the multi-modal feature attention mechanism (MMF) facilitates both inter-modality and intra-modality interactions by capturing the global correlations between modalities and employing dynamic sparse spatial attention. The effectiveness of MMF has been validated through comparative experiments, demonstrating its suitability for the phenotypic detection of artificially cultivated lettuce. In summary, the method proposed in this study facilitates real-time monitoring of facility crops, enabling precise control of environmental parameters in protected agriculture and optimizing resource allocation. This approach contributes to the development of a comprehensive intelligent agriculture system and establishes a foundation for unmanned farms.
Why it matches plant phenotyping methodsレタスの草丈、株幅、葉面積、 fresh weight、dry weightを推定するマルチモーダル画像ベース手法を開発し、実測値との比較およびアブレーション・比較実験で検証しており、フェノタイピング手法が研究の中心である。
abstracta specialized in situ phenotypic detection method has been developed
Citrus is widely loved for its rich nutritional value and unique flavor, and also occupies an important place in agriculture and the economy. The citrus industry has suffered severe losses in recent years due to the proliferation of citrus Huanglongbing (HLB). The transmission of HLB occurs via the Asian citrus psyllid insect vector and grafting practices, with no efficacious therapeutic intervention identified to date, aside from mitigating its dissemination through prompt identification and eradication of infected citrus trees. Detection of HLB is difficult due to its incubation period and the variety of symptoms at different stages of infection. Hence, there exists a pressing requirement for a detection methodology capable of integrating multi-level features of HLB to facilitate precise identification of the disease across various stages of infection. Most of the existing assays use single sensing, which leads to limitations and incompleteness in identifying specific markers induced by HLB. In this study, the effective configuration and complementarity of multimodal sensory information is achieved by establishing a fusion and complementary mechanism at the level of pre-processing and analyzing multisource information. The extraction of computer vision and electronic nose features of citrus leaves was realized using custom-developed portable detection devices. The performance of HLB detection was compared on different datasets obtained by multimodal feature fusion methods which include direct fusion method, stepwise fusion method and the improved Recursive Feature Elimination and Cross Validation (RFECV) feature selection method. The improved RFECV feature selection method uses the RFECV algorithm for each classification step in the delineated stepwise classification model and performs the feature set preference by cross-validation. The final improved RFECV feature selection method worked best for fusion of visual and olfactory features with an accuracy of 95.38% for HLB samples at various symptomatic stages, with 94.23% for early stage HLB and 94.12% for Zn Def. & HLB-positive samples. Multimodal feature fusion for feature acquisition proved to be superior to feature acquisition from a single sensing source, with enhanced fusion of HLB-induced feature sets at the visual and olfactory levels. It helps to improve the stability of the HLB measurement model to achieve the detection of HLB samples in complex environments. This method can provide generalized technical support for the application of multi-source sensing information and multimodal feature fusion methods in plant disease detection.
Why it matches plant phenotyping methods柑橘葉の症状を対象に、カスタム開発した画像処理・電子鼻装置とマルチモーダル特徴融合によるHLB検出法を開発・評価しており、植物病態の取得・推定が研究の中心である。
abstractThe extraction of computer vision and electronic nose features of citrus leaves was realized using custom-developed portable detection devices.
Accurate crop yield estimation is crucial for decision-making and planning in modern agriculture with increasing challenges with food security. Yield predictions provide farmers with insights into expected production, facilitating optimized resource allocation, improved agricultural management strategies, and enhanced profitability. This study investigates the application of machine learning (ML) techniques, including Feedforward Neural Networks (FNN), Long Short-Term Memory (LSTM), and Random Forest (RF) models, for predicting crop yields using multi-sensory time-series data that has been collected on two fields over a four-year timeframe. The focus is on corn (Zea mays) and cotton (Gossypium hirsutum) yield, two of the top critical crops in Mississippi region. A multi-sensory dataset was collected using multispectral cameras and LiDAR sensors mounted on unmanned aircraft systems (UAS), along with soil moisture and temperature data from volumetric probes and environmental data from a nearby weather station. Over four years, more than 30 features were extracted weekly from five major categories, with 235 ground truth yield records from plots in the field. The study outlines the methodology for feature selection and examines its impact on yield prediction accuracy. Using percentile root mean square error (RRMSE) and mean absolute percentage error (MAPE) as performance metrics, the study found that the proposed LSTM model produced lower field-wise errors (9 − 21 % MAPE) compared to other models and validation, indicating superior performance in predicting yields across selected weeks. The proposed ML-based approach, validated through year-based and field-wise cross-validation methods, demonstrates the effectiveness of using UAS-collected multi-sensor data for accurate yield estimation in corn and cotton.
Why it matches plant phenotyping methodsUASのマルチセンサー画像・LiDARデータから圃場プロットの収量を推定する計算ワークフローを開発・評価し、交差検証で性能を検証しているため、植物表現型取得が中心である。
abstractThis study investigates the application of machine learning (ML) techniques, including Feedforward Neural Networks (FNN), Long Short-Term Memory (LSTM), and Random Forest (RF) models, for predicting crop yields using multi-sensory time-series data
Efficient management and precise monitoring are essential for the sustainable control of crop diseases and pests. Traditional unimodal methods exhibit reduced reliability due to data gaps and environmental fluctuations. Multimodal artificial intelligence (AI) offers a promising alternative by integrating complementary data sources and enhancing robustness and adaptability. However, a comprehensive synthesis connecting multimodal AI with multi-scale disease and pest management is still lacking. Based on 950 publications from the past decade reflecting a 31.7% annual growth rate over the past five years, this review examines the evolution of AI-driven research and compares unimodal and multimodal approaches by summarizing major data modalities, fusion strategies, and modeling techniques. Deep learning emerges as the most widely used class of AI methods, and quantitative evidence indicates that multimodal systems achieve approximately 3–48.9% higher diagnostic accuracy than unimodal models. Evidence from 27 studies demonstrates the effectiveness of multimodal fusion across imaging, spectral, environmental, and sensor-based datasets. Building upon these findings, we propose a novel three-level management framework comprising point-level diagnosis, area-scale monitoring, and spatiotemporal forecasting, clarifying how multimodal AI strengthens each task. We further highlight the role of Plant Electronic Medical Records (PEMRs) and outline a conceptual virtual plant clinic to support continuous, data-driven crop health services. Finally, this review identifies key directions including advanced fusion strategies, lightweight and interpretable models, digital twin integration, and scalable decision-support systems, which are essential for intelligent and sustainable crop disease and pest management.
Why it matches plant phenotyping methods植物病害の診断・モニタリングを対象に、画像・スペクトル・環境・センサーデータを統合するAI手法とその性能を体系的にレビューしており、植物の病害状態推定が中心的な方法論的貢献である。
abstractthis review examines the evolution of AI-driven research and compares unimodal and multimodal approaches by summarizing major data modalities, fusion strategies, and modeling techniques.
This paper presents a three-phase deep learning framework comprising (i) multi-modal data acquisition from drones and satellites, (ii) standardized pre-processing including interpolation for missing temporal data, and (iii) CNN-based feature extraction for real-time health classification. This framework relies on a mathematical model based on neural networks that classifies and detects the condition of agriculture, removing the reliance on manual tasks and subjective diagnosis. This paper focuses on three main aspects of our framework: data acquisition, training and prediction. Data is collected using sensors like drones, cameras, and satellite imagery and is pre-processed to filter out noise and improve quality. The training part uses CNN to learn features from the data and become more meaningful. The prediction part of the task classifies, and diagnoses crop health through the trained model using the features. The framework accuracy for crops such as maize, potato, and wheat has been tested and yielded over 90% accuracy. The novelty of this work resides in the development of a multi-modal deep learning architecture that fuses macro-scale satellite imagery with micro-scale drone and IoT sensor data to improve diagnostic reliability. The framework was validated on a multi-source agricultural dataset using a 70% training, 15% validation, and 15% testing protocol. Experimental results demonstrate an accuracy exceeding 90% for staple crops. Using this framework can increase the visibility and quality of information maintained for crop health and improve the decision-making routine of farmers in real time. Additionally, automation of this process can significantly reduce labor costs and increase productivity per crop. Implementing this framework can contribute to precision agriculture and sustainable management practices.
Why it matches plant phenotyping methods作物の健康状態を植物の表現型・状態として推定するマルチモーダル画像・センサ基盤と深層学習手法を開発し、複数作物・データセットで検証しているため、方法が中心である。
abstractThis paper presents a three-phase deep learning framework comprising (i) multi-modal data acquisition from drones and satellites, (ii) standardized pre-processing including interpolation for missing temporal data, and (iii) CNN-based feature extraction for real-time health classification.
Reproduction assets foundThe article's Data availability section points to a public Kaggle dataset used for the crop classification/health diagnosis experiments, matching an allowed URL. No code or model checkpoints are disclosed.Dataset · publicript. The research work was guided by Dr. B.D.K.P. The Corresponding author Shshank Chaube collaborated for review and supervision. All authors reviewed the manuscript.
Funding
Open access funding provided by Symbiosis International (Deemed University). No funds, grants, or other support was received.
Data availability
Dataset: https://www.kaggle.com/datasets/bhagvendersingh/precision-agriculture-dataset .
Declarations
Competing interests
The authors declare no competing interests.
Ethical approval
This article does not contain any studies with human participants or animals performed by any of the authors.
References
1. Mohyuddin, G. et al. Evaluation of machine learning approaches for preciOpen asset ↗kaggle · bhagvendersingh/precision-agriculture-datasetlines:473-545Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
Rapid and accurate disease detection is essential for effective plant health management, which significantly increases crop yields and mitigates losses exacerbated by climate change. While machine-based imaging can detect symptoms invisible to the naked eye, standard models often overlook the influence of climatic and soil factors on disease occurrence. Objectives: To develop and assess the performance of deep learning models employing multimodal fusion strategies for the precise classification of crop diseases. Methodology: Proposed multimodal deep learning approach to combine visual features using MobileNetV2 with environmental and soil, including temperature, humidity, pH, rainfall, and nutrient levels (N, P, K) by the Multilayer Perceptron (MLP). Results: Experimental results demonstrate that MobileNetV2 excels among image-only models, proving to be an efficient lightweight architecture for agricultural tasks. Notably, the multimodal MobileNetV2 + MLP model achieved nearly 100% precision, significantly outperforming unimodal models. This highlights that incorporating environmental variables substantially enhances classification accuracy. Conclusion: Disease detection may become more precise if the multimodal correlations are strengthened by extending the dataset and upgrading the quality of environmental features, advanced fusion and alignment techniques usage that can further facilitate the interaction between the image and structured data, and lightweight CNNs, along with real-world testing that can help in the creation of reliable field-ready plant disease detection systems.
Why it matches plant phenotyping methods植物病徴を画像から分類するマルチモーダル深層学習手法の開発・性能評価が中心であり、植物の病害状態を直接推定している。
abstractObjectives: To develop and assess the performance of deep learning models employing multimodal fusion strategies for the precise classification of crop diseases.
Precise and timely prediction of wheat yield is pivotal for ensuring global food security and optimizing agricultural management practices, particularly through advanced remote sensing and machine learning techniques. In this study, wheat yield was accurately estimated by leveraging remote sensing-derived soil and vegetation indices. Yield data from 189 study points were collected, and Sentinel-2 (10-meter resolution) imagery from the Tillering and Anthesis growth stages was used. The models incorporated 25 variables, including 16 optical indices (e.g., NDVI, SAVI, MSAVI2) and three topographic factors. The key novelty of this research is the rigorous comparison of the predictive value and synergistic contribution of Sentinel-1 Synthetic Aperture Radar (SAR) data when integrated with Sentinel-2 optical indices within machine learning frameworks. Three machine learning approaches (Multiple Linear Regression (MLR), Support Vector Machine (SVM), and Random Forest (RF)) were employed and evaluated using 70% training and 30% testing subsets. Results revealed that the RF model, leveraging data from the Anthesis phenological stage, exhibited superior performance in wheat yield estimation, achieving an R² of 0.92 and an RMSE of 0.14 ton ha -1 for the training set, and an R² of 0.90 and an RMSE of 0.29 ton ha -1 for the testing set. To enhance model accuracy, Sentinel-1 radar data were integrated into the RF framework. This addition reduced the training set RMSE to 0.13 ton ha -1 but increased the testing set RMSE to 0.33 ton ha -1 , with R² values remaining stable at 0.92 and 0.90 for the training and testing sets, respectively. Variable importance analysis indicated that optical soil and vegetation indices were the dominant predictors. Although the inclusion of Sentinel-1 SAR data offered additional insights, it did not outperform the predictive capacity of optical indices. These findings validate the efficacy of the combined Sentinel-2 remote sensing approach for generating reliable wheat yield forecasts approximately 50 days prior to harvest.
Why it matches plant phenotyping methodsSentinel-1/2リモートセンシングと機械学習による小麦収量という植物形質の推定が中心で、モデル比較・評価と予測精度検証を実施している。
abstractwheat yield was accurately estimated by leveraging remote sensing-derived soil and vegetation indices.
Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 15 Sept 2026
ABSTRACT Sustainable agriculture urgently requires innovative, pesticide‐free strategies to mitigate herbivory and safeguard food security. Ultraviolet‐B (UV‐B) irradiation, with tunable intensity and cost‐effectiveness, has emerged as a promising non‐chemical method to enhance plant resistance, yet its underlying mechanisms remain elusive. Here, using tea plant ( Camellia sinensis ) and its major pest Ectropis obliqua as a model, we developed a multimodal framework that integrates AI‐enhanced electronic nose technology for real‐time volatile profiling with in situ hyperspectral stimulated Raman scattering (SRS) microscopy to characterize defense responses under precisely controlled UV‐B treatments. This approach identified herbivore‐induced volatiles—hexanal, (Z)‐3‐hexenol, octanal, and (Z)‐3‐hexenyl acetate—optimally induced at 1.2 kJ·m −2 UV‐B and linked to insect deterrence. SRS imaging further revealed elevated jasmonic acid derivatives and L‐phenylalanine, coupled with reduced protein levels and altered stomatal dynamics, all correlating with enhanced resistance. Transcriptomic and molecular analyses confirmed transcriptional regulation of these pathways. By bridging volatile detection, metabolic imaging, and molecular validation, this study pioneers a multimodal strategy that provides mechanistic insights into UV‐B–mediated plant defense and highlights the potential of multimodal methodologies as powerful tools for developing sustainable, pesticide‐free pest management solutions in precision agriculture.
Why it matches plant phenotyping methodsAI強化電子鼻とハイパースペクトルSRS顕微鏡を統合した植物防御応答のリアルタイム・多モーダル計測フレームワークが研究の中心であり、揮発性物質、代謝物、気孔動態などの植物状態を抽出している。
abstractwe developed a multimodal framework that integrates AI‐enhanced electronic nose technology for real‐time volatile profiling with in situ hyperspectral stimulated Raman scattering (SRS) microscopy to characterize defense responses
In the new age of greenhouse management, smart space adaptive systems are required to allow environmental control and crop analytics. This study presents an end-to-end deep learning architecture for spatiotemporal prediction and control in agriculture, which enables seamless integration between the different components of a decision loop. A ConvLSTM model predicts zone-specific microclimate variations, and a Dueling DQN agent selects the optimal actions for irrigation, ventilation, and fertilization according to energy demand, emission prediction, and soil moisture balance. The proposed EMMYP-Net (Enhanced Multimodal Multitask Yield Prediction Network) involves a CNN-BiLSTM-attention architecture that combines visual canopy data with multi-sensor sequences to co-classify growth stages and estimate yield. Experimental tests in a 1,200 m² four-zone greenhouse showed remarkable improvements, as ConvLSTM decreased RMSE by 45±2.3% compared to ARIMA and by 31±1.8% compared to LSTM. EMMYP-Net achieved an accuracy of 96.0% for classification, as well as an R² of 0.912±0.007 in predicting yields. This process-integrated approach enhanced resource sustainability by achieving savings of 19.3% in energy, 16.5% in water, and 15.2% in fertilizers relative to a conventional system. Combining predictive control and crop intelligence offers a scalable basis for sustainable data-driven greenhouse management. The key novelty of this work lies in the seamless integration of ConvLSTM-based spatiotemporal forecasting with the EMMYP-Net multimodal crop analytics within a unified reinforcement-learning-driven decision loop, enabling both predictive control and biological feedback in real time.
Why it matches plant phenotyping methods視覚的キャノピー情報とマルチセンサー系列から生育段階と収量を推定するEMMYP-Netが技術的中核であり、植物の状態・形質の計測手法として評価されている。温室制御全体を扱うが、フェノタイピング部分も実質的に記述・検証されている。
abstractThe proposed EMMYP-Net (Enhanced Multimodal Multitask Yield Prediction Network) involves a CNN-BiLSTM-attention architecture that combines visual canopy data with multi-sensor sequences to co-classify growth stages and estimate yield.
Fine-grained 3D phenotypic analysis of rice plays a vital role in rice breeding and yield estimation. However, a comprehensive rice data acquisition and segmentation pipeline is still lacking. While Neural Radiance Fields (NeRF) have shown impressive results in crop-level 3D reconstruction, their high sensitivity to data volume and camera viewpoints often leads to reconstruction failures for rice. In addition, the large-scale rice point clouds, coupled with heavy occlusion and visual similarity among grains, pose significant challenges for fine-grained trait extraction. To address the challenge of reconstructing rice point clouds under low-quality data conditions, we propose a novel method named Multi-Scale NeRF(MSNeRF). This method incorporates a structure-detail collaborative reconstruction mechanism and a dynamic initialization density scheduling strategy. Furthermore, we introduce a multimodal and multitask rice dataset (MMR) as a benchmark resource for future research. For rice point cloud segmentation, we develop Vision Rice Knowledge Graph Network(VRKGNet), which comprises an image segmentation module, a projection module, and a point cloud segmentation module enhanced with a Transformer to enlarge the receptive field. VRKGNet performs standalone point cloud segmentation and integrates image segmentation results from multiple viewpoints as prior knowledge to enhance semantic and instance-level segmentation. Extensive experiments demonstrate that MSNeRF achieves high-fidelity point cloud reconstruction with as few as 10 viewpoints. VRKGNet achieves superior rice plant segmentation with a semantic segmentation mIoU of 88.79% and an instance segmentation AP 25 of 84.55%, outperforming mainstream algorithms.
Why it matches plant phenotyping methods米の3D形質取得・再構成・分割を中核とする手法開発であり、データセット/ベンチマークも提供しているため、植物フェノタイピング手法文献に該当する。
abstractwe propose a novel method named Multi-Scale NeRF(MSNeRF)
Reproduction assets foundThe paper's authors explicitly state that the source code for MSNeRF and VRKGNet is publicly available on GitHub with testing scripts and test cases to reproduce the main results. The MMR dataset itself is only available upon request from the corresponding author, so it does not qualify as a public asset.Code · publicof Hefei Artificial Intelligence Breeding Accelerator Co. Ltd. ( NB2024005-02 ).
Data availability
The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper:
•
MSNeRF : https://github.com/qfwysw/MSNeRF.git
•
VRKGNet : https://github.com/qfwysw/VRKGNet.git
The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author.
Declaration of competing interest
The authors declare that they have no known competing financial iOpen asset ↗https://github.com/qfwysw/MSNeRF.gitlines:620-663Code · publicator Co. Ltd. ( NB2024005-02 ).
Data availability
The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper:
•
MSNeRF : https://github.com/qfwysw/MSNeRF.git
•
VRKGNet : https://github.com/qfwysw/VRKGNet.git
The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could haveOpen asset ↗https://github.com/qfwysw/VRKGNet.gitlines:620-663Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · OpenAlex · checked 15 Sept 2026
Hyperspectral imaging (HSI) has emerged as a powerful tool for precision agriculture, enabling the non-destructive monitoring of crop biochemical and physiological traits. However, HSI alone lacks structural context, which limits its ability to accurately capture complex canopy architectures and organ-level traits. Integrating HSI with depth-sensing modalities such as Light Detection and Ranging (LiDAR), Red, Green, Blue, and Depth (RGB-D) cameras, and computational reconstruction technique such as photogrammetry enables the generation of three-dimensional hyperspectral point clouds, combining spectral richness with geometric fidelity. This multi-modal fusion enhances crop trait estimation, including biomass, leaf chlorophyll content, canopy height, leaf area, and stress indicators, while improving the robustness of phenotyping under occlusions, shadows, and varying illumination. Dimensionality reduction, feature selection, and machine learning approaches, including deep learning and explainable AI, are useful for handling high-dimensional hyperspectral data and extracting actionable agronomic insights. Moreover, the integration of thermal, radar, and Global Navigation Satellite System (GNSS) data further expands the capabilities of multi-modal sensing, enabling continuous, all-weather crop monitoring and accurate spatial referencing. Despite these advances, most studies to date focus on controlled environments, highlighting the need for field-based validation to ensure the reliability and scalability of HSI-depth fusion techniques. This review consolidates current knowledge on multi-modal hyperspectral and 3D crop reconstruction, highlighting methods, applications, and challenges, and outlines future directions for implementing high-throughput, real-time phenotyping and precision agriculture solutions.
Why it matches plant phenotyping methods植物形質推定のためのハイパースペクトル・深度センシング融合と3D再構成を中心に扱うレビューであり、フェノタイピング手法の方法論的整理が主題。
abstractThis review consolidates current knowledge on multi-modal hyperspectral and 3D crop reconstruction, highlighting methods, applications, and challenges
Accurate growth stage recognition is vital for optimising crop inputs and improving yield efficiency. However, single-modal methods often fail to distinguish phenologically adjacent stages due to canopy similarity, spectral saturation, or structural ambiguity, leading to mistimed agronomic actions and resource loss. Moreover, the high computational demands of complex deep neural networks hinder practical implementation. To address this challenge, this paper proposes a lightweight multimodal data fusion network for wheat growth stage recognition. Specifically, the RGB images, multispectral (MS) data, and digital surface model (DSM) acquired by Unmanned Aerial Vehicles (UAV), along with derived spectral vegetation index (VI), are used to capture multimodal canopy features, including colour, spectral reflectance, and spatial structure. Furthermore, the approach utilises MobileNetV3-Small, a lightweight convolutional neural network, as the backbone to construct the multimodal data fusion framework. This framework enables efficient feature extraction and integration, achieving precise and rapid wheat growth stage recognition with minimal computational overhead. The results demonstrate that, compared to single-modal models, the proposed multimodal fusion model significantly enhances growth stage recognition accuracy, achieving an accuracy of 99.57 %, a precision of 99.58 %, a recall of 99.57 %, and an F1 score of 99.57 %. Notably, it improves stage differentiation in critical transitions such as Booting to Heading, reducing field misclassification risks and supporting quick decision-making. Comparative analysis with MobileNetV3-Large, ResNet-18, MNASNet, EfficientNet-B0, and ConvNeXt-Tiny demonstrates that MobileNetV3-Small offers the best trade-off between accuracy and resource efficiency, with only 1.53 M parameters and an inference time of 6.03 ms on RTX 4090 and 25.11 ms on Jetson Orin NX. This efficiency enables real-time deployment on resource-constrained edge devices, such as onboard UAV processors or in-field embedded systems. Overall, this study effectively overcomes the challenges of recognising adjacent growth stages and computational constraints, offering a robust theoretical foundation and an efficient, accurate solution for wheat growth stage recognition.
Why it matches plant phenotyping methodsUAV画像・マルチスペクトル・DSMを用いて小麦の生育段階という植物状態を推定する融合手法を開発・評価しており、表現型取得・抽出が研究の中心である。
abstractthis paper proposes a lightweight multimodal data fusion network for wheat growth stage recognition.
High-throughput field phenotyping bridges genotype, environment, and phenotypic performance. Conventional plot-level approaches relying on manual surveys are labor-intensive, and error-prone and fail to capture variability among individual plants, limiting seed cotton yield estimation and genotype screening under natural conditions. To address these limitations, a complex framework was developed, integrating single-plant instance segmentation, multi-trait inversion, plot-level stability characterization, and yield estimation. The enhanced vision-model framework, TopoRefineSAM, combines YOLOv12 detection with SAM2 segmentation and incorporates adaptive enhancement and topological refinement modules, enabling efficient, robust, and cost-effective single-plant identification under weak annotation. Based on this segmentation, multi-source UAV imagery (RGB, multispectral, thermal infrared, and DSM) was used to build ensemble learning models for inversion of physiological and biomass traits. A Stability Index Group (SIG) translates inter-plant variability into plot-level stability features, improving interpretability and consistency in yield estimation and cultivar screening. Results demonstrated that TopoRefineSAM achieved high segmentation accuracy for single-plant extraction under complex field conditions. In multi-trait inversion, Gradient Boosting Decision Trees (GBDT) achieved the highest performance. Our results demonstrated strong consistency between multimodal features and measured traits. In yield estimation, incorporating the SIG substantially improved predictive performance across growth stages. In cultivar screening, the method achieved high agreement with field measurements, showing robust identification of top-performing cultivars. Collectively, the findings establish a scalable, cost-effective, and high-accuracy framework for field-based phenotypic analysis and yield estimation, providing both methodological innovations and practical support for precision breeding and large-scale crop improvement.
Why it matches plant phenotyping methods単一植物のセグメンテーション、マルチモーダル画像による形質推定、安定性指標、収量推定を統合した圃場フェノタイピング手法の開発が中心である。
abstractTo address these limitations, a complex framework was developed, integrating single-plant instance segmentation, multi-trait inversion, plot-level stability characterization, and yield estimation.
Aiming at the hysteresis problem of traditional contact-type mechanical throughput detection method during combine harvester operation, this study proposes a multimodal data-driven online throughput prediction method. By building a multimodal sensor system that integrates vehicle-mounted cameras, GPS, grain moisture content, and feeding auger power sensors, a throughput prediction framework based on wheat ear biomass characteristics was established: Firstly, the WEC-MVFF wheat ear online counting density map estimation model is designed, and MobileViT is used to build a feature extraction backbone network. The multi-scale fusion module and the centralized conversion module are combined to realize the collaborative extraction of shallow texture features and deep semantic features of dense wheat ears in the field. Secondly, divide the image into regions of interest and establish a throughput prediction model based on multimodal information of wheat ear number (image) − moisture content (sensor) − travel speed. At the same time, a detection model based on feeding auger power is constructed as a comparison benchmark. Field tests show that the WEC-MVFF model maintains an average counting accuracy of more than 90 % under different wheat ear density, travel speed and light intensity conditions. The model’s online counting advantage is verified through ablation study and comparative tests with other counting models. The throughput prediction method achieved an MAE of 0.70 kg/s and 0.77 kg/s in test area 1 and 2 respectively, the first-order difference fluctuation was stable in the range of ±0.50 kg/s, and the prediction frame rate of 10-13fps met the real-time requirements. Compared with the single-modal image prediction method, the accuracy of the multimodal method considering moisture content was improved by 0.80 kg/s and 0.75 kg/s in the two test areas, respectively. Compared with traditional mechanical quantity detection methods, it has higher accuracy and stability while achieving early prediction, providing reliable feedforward information support for the intelligent control of harvesters.
Why it matches plant phenotyping methods小麦穂の画像計数とマルチモーダルセンサを用いて、穂密度・バイオマス特性および収量流量を推定する手法を開発・検証しており、植物形質の取得・推定が研究の中心である。
Advancements in phenotyping technologies, including object imaging, high-throughput monitoring, and soft computing, are pivotal for understanding plant responses to environmental stresses. These technologies enable detailed analyses of morphological, physiological, and structural adaptations under abiotic and biotic stresses, such as drought. Current work using multimodal and multi-perspective image processing methods can capture the essential processes that enhance plant resilience and counteract stress by identifying morphological and biochemical indicators. However, the dynamic and complex nature of plant responses poses multiple challenges for generating precise analytics and descriptors of evolving phenotypes. This work introduces analytics for concurrent imaging, adopting the underlying principle of cosegmentation to create taxonomies for new phenotypes. Here, unidimensional refers to the concurrent analysis of multiple images within a single phenotyping dimension: temporal, modal, or perspective, rather than combining information across dimensions. The proposed unidimensional phenotypes integrate concurrent images within individual temporal, modal, or perspective dimensions to capture dynamic morphological and physiological responses that are not observable with conventional single-image or cumulative metrics. Within a high-throughput imagery production system, these phenotypes enable more nuanced quantification of phenotypic changes, leveraging the strengths of simultaneous image analysis to enhance insight into plant adaptations. This workflow aligns with the investigation of plants’ adaptive strategies under abiotic stress and provides quantitative indicators of plant health under adverse environmental conditions.
Why it matches plant phenotyping methods植物の同時画像解析とコセグメンテーションに基づく新しい表現型抽出・定量化ワークフローを提案しており、植物フェノタイピング手法が中心である。
abstractThis work introduces analytics for concurrent imaging, adopting the underlying principle of cosegmentation to create taxonomies for new phenotypes.
Reproduction assets foundThe paper's Data Availability Statement explicitly states that the SIMID and SIPID image datasets created and used in this study are publicly available on Zenodo (DOI 10.5281/zenodo.17400167), which is an allowed URL. These are the paper-specific plant phenotyping imagery inputs (buckwheat and sunflower under control/dDataset · publicThe SIMID and SIPID dataset utilized and created in this study is publicly available and accessible at the following link: https://doi.org/10.5281/zenodo.17400167Open asset ↗Zenodo · 10.5281/zenodo.17400167lines:174-216Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
Multimodal approaches for crop disease detection have gained significant attention due to their ability to integrate diverse data sources for improved accuracy. This review categorizes recent studies into five areas: multimodal deep learning and vision transformers, hyperspectral and remote sensing, thermal imaging and UAV applications, CNN–Transformer hybrids and ensemble methods, and comprehensive reviews. Results indicate that frameworks combining RGB, hyperspectral, and thermal imaging achieve accuracies up to 97.8%, while hybrid CNN–Transformer architectures reach 99.7% on benchmark datasets. Despite these advances, challenges remain in scalability, computational cost, and real-world deployment, highlighting the need for lightweight, explainable, and field-validated models.
Why it matches plant phenotyping methods植物病害を画像・リモートセンシングから推定する方法を体系的にレビューしており、病害状態という植物表現型の取得・推定が中心です。
abstractMultimodal approaches for crop disease detection have gained significant attention due to their ability to integrate diverse data sources for improved accuracy.
Accurate, near real-time soybean phenology information is critical for crop management and breeding. Previous approaches relying on satellite remote sensing time-series data suffer from temporal delays, limiting their usefulness for in-season decision-making. To overcome this limitation, this study reframes phenology identification as a near real-time classification task using single-timepoint Unmanned Aerial Vehicle (UAV) imagery collected from 420 soybean germplasm resources across three experimental sites, and proposes an innovative multi-modal dynamic Gating Fusion Model that integrates two optimized pathways. one based on machine learning (ML) and the other on deep learning (DP). In the ML branch, systematic benchmarking of tabular-feature models identified the Soft Voting ensemble as the best classifier. In the DL branch, an enhanced BC-ConvNeXt model equipped with BiFPN and CBAM modules was developed to strengthen visual feature extraction. Building on these two optimal classifiers, the dynamic gating fusion model achieved the highest F1-score of 94.3% across seven key growth stages (V1, V2, R1, R2, R6, R7, R8). This result represents a significant improvement of 1.5% and 10.6% over the best performing ML and DL models, respectively. The superior performance arises from the intelligent arbitration of complementary strengths, with gating-weight analysis revealing a strategy that prioritizes ML predictions while leveraging DL for error correction. This work establishes a complete framework for near real-time crop phenology detection and demonstrates the strong potential of intelligent multi-modal fusion in high-throughput phenotyping.
Why it matches plant phenotyping methodsUAV画像からダイズの生育段階を推定するマルチモーダル分類モデルを開発・ベンチマークしており、植物表現型取得手法が研究の中心である。
abstractproposes an innovative multi-modal dynamic Gating Fusion Model
• Very-high-resolution UAVs dominate row/plant analyses; Sentinel-2 underpins regional monitoring. • Multi-sensor fusion (UAV/satellite) and 3D bring robustness to vine identification. • Deep Learning enables detection of plots and accurate row delineation in complex terrains. • Deep Learning models requires vast annotated data and high computational cost limits routine use. • Validation and model portability remain the weakest methodological areas. Sustainable vineyard management and planning require reliable methods for identification and monitoring. This systematic review synthesises and appraises the literature on automatic vineyard identification using remote sensing (RS), from classical techniques to artificial intelligence (AI), describing the state of the art, patterns, challenges, and gaps. Guided by PRISMA and informed by selected SWiM reporting items, we conducted a systematic search across multiple databases, gathering all relevant records up to 13 July 2025, and included 108 sources, of which 80 empirical studies contributed to the synthesis. The risk of bias was assessed by adapting the principles of PROBAST-AI and QUADAS-2 to the agricultural context, covering data representativeness, sensors/pre-processing, ground-truth, validation, and portability; its application also guided the selection and organisation of the synthesis. The analysis was narrative and structured by scale and application objective (regional, parcel, row, and plant). The most common tasks were classification (28%), detection (26%), and segmentation (24%), with multitask pipelines being frequent. We observe a clear transition from pixel-based approaches using satellite imagery to methodologies that integrate very-high - resolution UAV imagery, 3D reconstruction, and Deep Learning (DL). UAVs dominate row and plant-level analyses, whereas Sentinel-2 has become the main tool for multitemporal regional monitoring. DL models, such as CNNs and Vision Transformers (ViTs), tend to deliver superior performance in canopy segmentation and parcel classification. The assessment identified model validation as the weakest methodological domain across studies. The main limitations lie in weak spatiotemporal portability of models and high computational costs, aggravated by reliance on large volumes of annotated data. Promising directions include multisensory fusion (UAV + satellite) and the integration of 3D information into DL pipelines, which increase robustness and operational applicability. These advances are enabling high-value, specialised objectives such as mapping in complex terrain, detecting abandoned vineyards, and identifying missing plants.
Why it matches plant phenotyping methodsブドウ園・列・植物レベルの自動識別とモニタリングに用いるリモートセンシング手法を体系的にレビューし、センサー、3D再構成、深層学習、検証性、可搬性を評価しており、植物状態・構造の取得手法が中心である。
abstractThis systematic review synthesises and appraises the literature on automatic vineyard identification using remote sensing (RS), from classical techniques to artificial intelligence (AI), describing the state of the art, patterns, challenges, and gaps.
Abstract Purpose Taro (Colocasia esculenta (L)) , a neglected and underutilized crop species (NUS), holds great potential as a future smart crop that can thrive under climate variability and change, hence sustaining food security. While taro exhibits tolerance to drought conditions, variations in physiological attributes such as leaf temperature that rises under water stress and the associated stomatal closure that is initiated to conserve water, compromise crop productivity and overall yield. Therefore, monitoring taro crop physiological indicators of water status allows for the implementation of timely interventions and targeted adaption strategies to mitigate the effects of water deficit on taro crop productivity. Methods Unmanned Aerial Vehicles (UAV), integrated with high-resolution thermal sensors, provide valuable platform for generating near-real-time spatially explicit information suitable for assessing taro crop water status physiological indicators at farm scale. Hence, this study sought to evaluate the utility of UAV multi-modal thermal remote sensing and deep neural network techniques to estimate the equivalent water thickness, fuel moisture content, stomatal conductance, canopy temperature, and the chlorophyll content of smallholder taro crops. Results Findings showed that the multi-modal variable method achieves higher estimation accuracies in comparison to a single-modal technique, achieving R 2 values greater than 0.91 and rRSME values less than 14.15% of equivalent water thickness, fuel moisture content, stomatal conductance, canopy temperature, and chlorophyll content. Additionally, the results illustrated that the thermal wavebands and derived thermal indices are the most influential variables in estimating stomatal conductance and leaf temperature, yielding R 2 of 0.96 and 0.95, respectively. Conclusion These research findings underscore the applicability of UAV-acquired thermal remote sensing in providing rapid and robust spatially explicit information on smallholder taro crop water status for ensuring crop productivity and developing early warning systems of water stress. These findings serve as a stepping stone towards advancing agricultural monitoring frameworks and integrating NUS, such as taro, into traditional farming.
Why it matches plant phenotyping methodsUAV熱・マルチスペクトルデータと深層学習により、タロイモの水分状態、生理形質、クロロフィルなどを推定し、精度も評価しているため、植物フェノタイピング手法の応用・技術評価が中心である。
abstractthis study sought to evaluate the utility of UAV multi-modal thermal remote sensing and deep neural network techniques to estimate the equivalent water thickness, fuel moisture content, stomatal conductance, canopy temperature, and the chlorophyll content of smallholder taro crops.
Introduction Ginseng, as a precious medicinal plant, requires precise classification of its seeds, which directly impacts production processes and the stability of herbal quality. Furthermore, this classification plays a critical role in advancing ginseng breeding and the modernization of the industry. Current research indicates that systematic automated precision classification technologies for ginseng seeds remain underdeveloped, necessitating breakthroughs in technical bottlenecks. Methods This study innovatively proposes a smart classification method based on multimodal data fusion. It employs recursive feature elimination (RFE) to select morphological features from images, followed by competitive adaptive reweighted sampling (CARS) to extract spectral bands from hyperspectral data within the 350~2500 nm range. Morphological and spectral features are then integrated to construct a random forest (RF) classification model optimized using an enhanced, red-billed blue magpie optimization (RBMO) algorithm. To address the RBMO algorithm's tendency to converge to local optima, the hybrid optimization framework is constructed by integrating three mechanisms: the improved Circle chaotic map, the golden sine search strategy, and the adaptive simulated annealing perturbation mechanism. Results Experimental results demonstrate that the proposed model outperforms the baseline model RF, achieving 4.69%、4.79%、4.69 and 4.74% improvements in classification accuracy, precision, recall, and F1-score on test datasets, respectively. Discussion The established multimodal data fusion classification system not only provides theoretical and technical foundations for industrial-scale ginseng seed classification but also offers a transferable intelligent decision-making paradigm for non-destructive testing in traditional Chinese medicine.
Why it matches plant phenotyping methods画像由来の形態特徴とハイパースペクトル特徴を統合し、種子を非破壊分類する手法自体が研究の中心であり、植物器官(種子)の観測可能な状態を推定しているため。
abstractThis study innovatively proposes a smart classification method based on multimodal data fusion.
Sustainable agriculture in arid regions faces critical challenges due to water scarcity, high temperatures, and inefficient traditional farming practices. This study presents an AI-enabled smart farming framework for optimizing date palm (Phoenix dactylifera) cultivation through the integration of Machine Learning (ML) and Internet of Things (IoT) technologies. A structured multimodal dataset comprising biometric features palm height, trunk diameter, and leaf number, environmental parameters soil moisture, temperature, and humidity, and categorical attributes variety and health status was analyzed to classify palm health and support data-driven irrigation management. Four ML algorithms Random Forest (RF), Gradient Boosting Machine (GBM), Artificial Neural Network (ANN), and Support Vector Machine (SVM) were developed and optimized using grid search with five-fold cross-validation. Among them, the Random Forest model achieved the highest classification accuracy of 95.3%, demonstrating strong robustness for heterogeneous agricultural data. Feature importance analysis highlighted soil moisture, humidity, trunk diameter, and leaf number as key contributors to palm health prediction. The proposed AI-IoT framework enables real-time monitoring, predictive diagnostics, and automated decision support for sustainable water use and crop management, aligning with Saudi Vision 2030 objectives for technology-driven and resource-efficient agriculture.
Why it matches plant phenotyping methodsヤシの生体特徴から健康状態を分類する機械学習手法を開発・比較し、分類性能を評価しているため、植物状態推定が中心的な方法的貢献である。
abstractThis study presents an AI-enabled smart farming framework for optimizing date palm (Phoenix dactylifera) cultivation through the integration of Machine Learning (ML) and Internet of Things (IoT) technologies.
Cassava remains a critical food-security crop across Africa and Southeast Asia but is highly vulnerable to diseases such as cassava mosaic disease (CMD) and cassava brown streak disease (CBSD). Traditional diagnostic approaches are slow, labor-intensive, and inconsistent under field conditions. This review synthesizes current advances in combining unmanned aerial vehicles (UAVs) with deep learning (DL) to enable scalable, data-driven cassava disease detection. It examines UAV platforms, sensor technologies, flight protocols, image preprocessing pipelines, DL architectures, and existing datasets, and it evaluates how these components interact within UAV–DL disease-monitoring frameworks. The review also compares model performance across convolutional neural network-based and Transformer-based architectures, highlighting metrics such as accuracy, recall, F1-score, inference speed, and deployment feasibility. Persistent challenges—such as limited UAV-acquired datasets, annotation inconsistencies, geographic model bias, and inadequate real-time deployment—are identified and discussed. Finally, the paper proposes a structured research agenda including lightweight edge-deployable models, UAV-ready benchmarking protocols, and multimodal data fusion. This review provides a consolidated reference for researchers and practitioners seeking to develop practical and scalable cassava-disease detection systems.
Why it matches plant phenotyping methodsUAV画像と深層学習によるカッサバ病害の検出手法を中心に、センサー、撮影プロトコル、画像処理、モデル、データセット、性能指標を体系的にレビューしているため、植物の病害状態を対象とするフェノタイピング手法レビューに該当する。
abstractThis review synthesizes current advances in combining unmanned aerial vehicles (UAVs) with deep learning (DL) to enable scalable, data-driven cassava disease detection.
The purpose of Precision Agriculture is to incorporate technology into various agricultural processes to increase efficiency and productivity. In fact, Precision Agriculture uses advanced technologies such as sensors and data analytics to improve crop yields. However, a significant challenge in this area is effectively integrating multiple data sources to accurately predict crop health and yield using all available information. This problem arises because traditional models typically use spectral analysis or deep learning techniques independently. Due to this separation, neither method generates the desired results. Researchers propose a solution to this issue by combining spectral analysis and deep learning for multimodal data fusion in precision agriculture. Our integrated approach begins with the collection of multispectral data from drone- or satellite-based sensors to characterise crop types. Spectral analysis will determine each crop type's chlorophyll and water content, which affect plant health. Deep learning will be used to analyse the intricate interconnections between crop yields and derived attributes to understand their relationships better. Integrated use of these two technologies will give us a broader range of data and knowledge about crop variety health and yield than single-use or standalone applications.
Why it matches plant phenotyping methodsマルチスペクトルセンサーと深層学習を統合し、作物のクロロフィル・水分量、健康状態、収量を推定する方法が研究の中心である。
titleIntegrating Deep Learning and Spectral Analysis for Multi-Modal Data Fusion in Precision Agriculture for Enhancing Crop Health Monitoring and Yield Prediction
Crop diseases and pests pose significant threats to global food security, demanding precise and efficient management solutions. While Multimodal Large Language Models (M-LLMs) offer promising avenues for intelligent agricultural diagnosis, general-purpose models often falter due to a lack of specialized visual feature extraction, inadequate understanding of agricultural terminology, and insufficient precision in prevention advice. To address these challenges, this paper introduces AgriM-LLM, a novel agriculture-specific multimodal large language model designed for enhanced crop disease and pest identification and prevention. AgriM-LLM integrates several key innovations: an Enhanced Vision Encoder featuring a Multi-Scale Feature Fusion module for capturing subtle visual symptoms; an Agriculture-Knowledge-Enhanced Q-Former that injects structured agricultural knowledge to guide cross-modal alignment; and a Domain-Adaptive Language Model employing a multi-stage progressive fine-tuning strategy for expert-level advice generation. Furthermore, an efficient LoRA-based fine-tuning strategy ensures practical computational resource utilization. Evaluated on a comprehensive Chinese agricultural multimodal dataset, AgriM-LLM consistently outperforms existing general-purpose and domain-specific baselines. Our ablation studies confirm the critical contribution of each proposed component, and detailed analyses demonstrate superior visual encoding, knowledge integration, and linguistic specialization. AgriM-LLM represents a significant step towards providing timely, accurate, and actionable intelligent decision support for farmers, thereby fostering sustainable agricultural development.
Why it matches plant phenotyping methods作物の視覚症状から病害を識別するマルチモーダルモデルを開発・評価しており、植物の病害状態の推定が中心的な技術貢献です。ただし害虫管理や農業意思決定支援も含みます。
abstractthis paper introduces AgriM-LLM, a novel agriculture-specific multimodal large language model designed for enhanced crop disease and pest identification and prevention.
Seed phenomics is a critical research field for understanding seed germination mechanisms. Metasurfaces, composed of subwavelength nanostructures, offer a promising pathway to achieve both dispersion control and imaging functionalities within an ultra-compact form factor. Recent advances in micro–nano-optics and computational imaging have opened new avenues for high-dimensional, multimodal imaging. However, conventional hyperspectral and light-field systems still face limitations in compactness, depth resolution, and spectral–spatial integration. This review summarizes recent progress in metalens and metasurface lens array-based light-field systems for hyperspectral imaging and 3D reconstruction, with a focus on the underlying principles, design strategies, and reconstruction algorithms that enable single-shot 3D hyperspectral acquisition. We further present a forward-looking roadmap toward the realization of a revolutionized imaging paradigm: a metasurface-based light-field platform that fully integrates 3D and hyperspectral imaging capabilities. In particular, we examine how dispersive metasurfaces serve as core optical elements for precise dispersion control in hyperspectral imaging systems, while metalens arrays enable accurate modulation of spatial–angular distributions in light-field configurations. We systematically review both 3D and spectral reconstruction algorithms, highlighting their roles in decoding complex optical encodings. The application of these integrated systems in seed phenotyping is emphasized, demonstrating their capability to capture 3D spatial–spectral distributions in a single exposure. This approach facilitates high-throughput analysis of morphological traits, germination potential, and internal biochemical composition, offering a comprehensive solution for advanced seed characterization. Finally, we outline a practical roadmap for implementing a metasurface-based light-field platform that integrates hyperspectral imaging and computational 3D reconstruction. This review offers a comprehensive overview of the state of the art in compact 3D light-field systems and multimodal hyperspectral imaging platforms, while providing forward-looking insights aimed at advancing smart seed phenotyping, precision agriculture, and next-generation optical imaging technologies.
Why it matches plant phenotyping methods種子フェノミクス向けの3D・ハイパースペクトル画像取得および再構成プラットフォームを中心にレビューし、形態形質や発芽能の解析への適用を扱うため、方法論が中心です。
abstractThis review summarizes recent progress in metalens and metasurface lens array-based light-field systems for hyperspectral imaging and 3D reconstruction
Early warning of overgrowth in strawberry seedlings is essential to balance vegetative and reproductive growth. However, existing monitoring methods face major challenges, including subtle visual symptoms and limited abnormal samples. To address this, we propose MM-CAPNet, a multimodal fusion framework for early detection of seedling overgrowth. We first developed a representative sample collection of strawberry seedlings through a systematic induction experiment, integrating historical environmental time-series data with contemporaneous plant images. The MM-CAPNet architecture uses a dual-stream design to process these inputs, with a Transformer encoder for environmental sequences and a MobileNetV2 encoder for images. A critical component of the proposed framework lies in the image-guided Cross-Attention mechanism, which uniquely treats the current phenotype as an active query to adaptively retrieve and aggregate the most diagnostically relevant segments of past environmental data. Experiments show MM-CAPNet outperforms baselines, reaching 87.6% accuracy and 0.901 AUC, with strong discriminative ability for early overgrowth categories. Ablation studies confirm its interpretability by linking visual phenotypes to key environmental drivers. This work provides growers with a proof-of-concept framework to regulate fertilization, irrigation, and light management during the nursery stage, thereby reducing the risk of excessive vegetative growth. The proposed framework supports precision cultivation strategies that enhance resource efficiency and crop resilience.
Why it matches plant phenotyping methods画像と環境時系列を統合し、イチゴ苗の過繁茂という植物状態を早期推定する新規マルチモーダル手法を開発・評価しており、表現型取得・判定が研究の中心である。
abstractwe propose MM-CAPNet, a multimodal fusion framework for early detection of seedling overgrowth.
Moisture plays a critical role in crop growth and development, making accurate, efficient, and non-destructive detection and monitoring of crop water stress essential for advancing crop science research and optimizing production management. Traditional non-destructive methods for monitoring water stress primarily rely on color imaging or partial 2D spectral analysis. However, these methods are limited to two-dimensional features and fail to capture the spatial variability of water stress within the three-dimensional canopy structure of crops. To address this limitation, this study integrates RGB-D cameras and thermal infrared cameras and introduces a method for calculating the 3D spatial distribution characteristics of crop water stress using RGB-D-T fusion analysis. This approach enables high-precision detection and analysis of water stress in strawberry plants. An RGB-D-T acquisition system was designed and implemented to collect RGB images, depth images, and thermal infrared images of strawberries subjected to different moisture gradient treatments. Using the YOLOv8-seg deep learning model, semantic segmentation of the crop canopy and the wet reference surface was performed. The segmentation results were fused with 3D point cloud data to generate a 3D dataset incorporating temperature, color, and semantic information. Subsequently, the three-dimensional distribution characteristics and dynamic changes in the canopy water stress index (CWSI) of strawberry plants were analyzed under varying moisture conditions. The results demonstrated that under low moisture gradients (15%–30%), the CWSI value increased significantly and exhibited a concentrated distribution, indicating severe water stress. Conversely, under high moisture gradients (75%–90%), the CWSI value approached zero, reflecting sufficient water supply and complete stress alleviation. Additionally, the study highlighted the variation in the temperature difference between strawberry leaves and the surrounding air, confirming the sensitivity of strawberries to water stress across different reproductive stages. The response to water deficit was most pronounced during the growth phase. By fusing multi-source data, this study achieves 3D visualization and precise quantification of water stress in strawberries, providing innovative insights and technical support for precision irrigation and crop phenotyping research.
Why it matches plant phenotyping methodsRGB-D・熱赤外センサーの融合、3D点群化、深層学習セグメンテーションにより、イチゴの水ストレスを3D定量化する取得・解析手法が研究の中心である。
abstractAn RGB-D-T acquisition system was designed and implemented to collect RGB images, depth images, and thermal infrared images of strawberries subjected to different moisture gradient treatments.
Efficient and accurate extraction of plant height (PH) plays an important role in analyzing its deeper phenotypic traits and improving breeding efficiency. Traditional methods make it difficult to measure PH at the individual plant level in the plot-level on a large scale, with high accuracy and low delay. To address this issue, we adopted a truss-type phenotyping platform to acquire LiDAR and RGB canopy data from 120 rapeseed genotypes from the two-leaf stage to the flowering stage, covering a total of 6 growth stages. (1) Object detection was performed to identify per plant of rapeseed images by the K-Means improved Faster R-CNN algorithm. (2) The image data before and after object detection and the point cloud data were fused to recognize per plant on the point cloud. Besides, the rapeseed plant point cloud and the ground point cloud were distinguished by color. (3) The Cloth Simulation Filter (CSF) algorithm is used to fit the ground points obscured by the canopy, which contributes to accurately extracting the PH of individual rapeseed. The field tests indicated that the mAP (IoU = 0.5) of the improved object detection method was 0.902, which achieved a high detection accuracy for rapeseed with different sizes in all periods. Specifically, in the PH accuracy verification of the No. 077 cultivar, R² was 0.997, RMSE was 1.156 cm, rRMSE was 5.48 %, and the maximum difference between automatic recognition and manual measurement in the late stage of growth was less than 6 cm. Especially, the proposed method can extract of the height of individual plant with the values of R² was 0.980, RMSE was 0.651 cm, rRMSE was 6.543 % in the seedling stage and the PH was lower than 20 cm. The differences of pH of 40 rapeseed genotypesunder cold stress were compared, which provide a reference for high-throughput PH extraction and exploring genotypic differences in plant breeding.
Why it matches plant phenotyping methodsLiDAR・RGBデータ融合と画像/点群処理により、個体ごとの草丈を高精度・ハイスループットに抽出する手法の開発と検証が中心であるため。
titleHigh-throughput extraction of individual plant height in rapeseed based on LiDAR-Camera data fusion
Published1 Jan 2026Journal of food composition and analysis : an official publication of the United Nations University, International Network of Food Data Systems
Soluble solids content (SSC) is an important indicator for determining the commercial value of peaches. Visible/near-infrared (Vis/NIR) spectroscopy combined with chemometric methods is a primary technique for predicting peach SSC. However, the interference of fruit color with spectral signals makes it challenging to accurately detect SSC across different varieties. This study explored the feasibility of fusing spectral and image data to achieve accurate SSC prediction for multiple peach varieties. Diffuse reflectance spectra and images of three peach varieties (‘Hujing’, ‘Jinqiuhong’, and ‘Dongxue’) were collected. Multiple feature-level fusion strategies for spectral and image data were proposed. Partial least squares regression (PLSR) and support vector regression (SVR) models were developed based on the multimodal fusion data to predict the SSC of individual and multiple varieties, respectively. Their predictive performance was compared with that of models established using spectral data alone. To further improve the generalization ability of the multi-variety models, a spectrum-image fusion network (SIFNet) was proposed by extracting and leveraging high-level image features and integrating them with spectral information. The results showed that the SIFNet achieved superior performance in predicting the SSC of multi-variety peaches, with RP2, RMSEP, and RPDP of 0.8235, 1.0514, and 2.5208, respectively.
Why it matches plant phenotyping methodsスペクトル・画像融合とSIFNetを開発し、個々のモモ果実のSSCという器官形質を予測する方法が研究の中心である。
abstractThis study explored the feasibility of fusing spectral and image data to achieve accurate SSC prediction for multiple peach varieties.
Tomato, as a globally important economic crop, requires precise and timely disease management to secure yield and quality. Yet segmentation robustness is often limited by weak semantic understanding from single-modality images, narrow receptive fields of convolutional structures, and discontinuous boundary predictions. To address these issues, we propose the Multi-scale Linear Cross-modal Fusion Architecture for Tomato Leaf Disease Segmentation (MS-LCFNet). We construct a real-world field dataset covering five major tomato leaf diseases, annotated by experts with detailed textual descriptions to enable multimodal learning. MS-LCFNet strengthens semantic representation via cross-modal fusion, captures local and global context through an Adaptive Long-short Distance Perception module, and improves boundary continuity with a Physics-informed Smoothness-constrained Loss. Experiments show that MS-LCFNet achieves 87.13 % mIoU on our dataset and 90.78 % on PlantVillage, improving over previous state-of-the-art methods by + 4.62 % and + 4.48 %, respectively, and demonstrating superior accuracy and robustness in complex agricultural scenarios.
Why it matches plant phenotyping methodsトマト葉の病害状態を画像からセグメンテーションする手法を開発し、独自データセットで性能評価しており、植物病害表現型の取得・抽出が中心である。
titleA multi-scale linear cross-modal fusion architecture for tomato leaf disease segmentation
Controlled-environment agriculture (CEA) and circular production systems require coordinated monitoring of biological and physicochemical processes across trophic levels. This project report presents the implementation of a multi-trophic controlled-environment agriculture demonstrator that integrates computer-vision-based monitoring with established sensor infrastructure for aquaculture, poultry, plants, microalgae, duckweed, and insect modules. Stereo imaging and RGB-D systems are deployed for non-invasive quantification of fish biomass and plant growth, while continuous water-quality and environmental measurements (e.g., pH, dissolved oxygen, nitrate, ammonium, temperature, CO$_2$) provide complementary process data. These data streams are synchronized within a shared database architecture to enable cross-module evaluation of nutrient dynamics, growth progression, and operational stability under real facility conditions. The implemented framework demonstrates how computer vision can extend conventional sensor-based monitoring by directly capturing biological performance indicators across aquatic, terrestrial, and microbial domains. While advanced predictive modeling and full digital twin simulation remain future development steps, the realized data-integration architecture establishes a structural foundation for the systematic evaluation of circular indoor food-production systems. The demonstrator illustrates how multimodal monitoring can support nutrient recirculation, transparency of biological variability, and data-driven assessment within controlled multi-trophic environments.
Why it matches plant phenotyping methods植物成長をステレオ画像およびRGB-Dで非侵襲的に定量するコンピュータビジョン監視基盤を実装しており、植物フェノタイピングが統合監視システムの主要な技術要素である。
abstractStereo imaging and RGB-D systems are deployed for non-invasive quantification of fish biomass and plant growth
Agricultural systems are entering an era defined not just by mechanization but by real-time, spatially aware, data driven production practices. This shift which has rapidly evolved from novelty to necessity in precision agriculture. At the center of this shift Unoccupied Aerial Vehicles (UAV) or drones have operationalized high-resolution aerial data into sustainable agricultural outcomes. This chapter provides an application-focused roadmap for integrating UAV into modern agronomic workflows as pre, during, and post flight operations for agronomic flight planning. It bridges the engineering of flight platforms and sensors with the applications in crop stress detection, variable-rate input application, and predictive yield modeling. We explore UAV architecture and their implications for data resolution, field scale, and operational complexity. Sensor systems use the portions of electromagnetic spectrum (RGB, multispectral, hyperspectral, thermal, LiDAR) and are examined through their applications for plant physiology and soil interactions. Ground sampling distance, spectral calibration, geospatial accuracy and photogrammetry are not treated as ancillary steps, but as critical determinants of agronomic utility. We describe the full data pipeline from FAA (Federal Aviation Administration) regulations to machine learning-driven analytics, acquisition, orthomosaic generation, digital surface modeling, vegetation index extraction, and the development of actionable prescription maps. In the context of AI evolution, we emphasize how AI-ML methods classification, regression, clustering, and dimensionality reduction help with integrating complex patterns in time-series UAV imagery, enabling early and precise management of nutrients, water, weeds, and disease. By synthesizing global regulatory frameworks and field-based use cases, the chapter concludes UAV as tools that transforms data into agronomic decisions.
Why it matches plant phenotyping methodsUAVセンサー、校正、フォトグラメトリ、オルソモザイク、植生指数抽出、機械学習解析を含む一連の植物状態・ストレス推定ワークフローをレビューしており、単なる生物学的実験の測定ではなく、取得・解析手法とプラットフォームが中心です。
abstractThis chapter provides an application-focused roadmap for integrating UAV into modern agronomic workflows as pre, during, and post flight operations for agronomic flight planning.
Sap flow serves as the primary carrier for water, nutrients, and signaling molecules, playing a crucial role in fruit development by delivering these essential constituents to the fruit. While the efflux of sap from fruit to other organs (termed reverse sap flow) has been observed in plants, its underlying mechanisms remain unclear due to a lack of effective methodologies for comprehensive studies. Here, we pioneered the integration of real-time sap flow measurements from novel plant-wearable sensors with synchronized environmental monitoring, establishing a multimodal data framework to systematically decode the endogenous causes and exogenous triggers of reverse sap flow in watermelon plants. Our experimental results reveal that plant water supply-consumption imbalance is the core endogenous cause of reverse sap flow, which is induced by two external triggers in the natural environment: rapid light intensity surges and soil drought. Furthermore, a long-term drought stress experiment illustrates that reverse sap flow from the fruit enhances the drought resistance of plants by adjusting water redistribution within the whole plant. This study challenges the unitary view of fruit solely as a "sink" in the traditional source-sink theory, further refines the understanding of the source-sink paradigm, and provides a novel mechanism and insight for plant drought tolerance strategies.
Why it matches plant phenotyping methods新規の植物ウェアラブルセンサーによるリアルタイム樹液流計測と環境モニタリングの統合が研究の中心で、植物の水輸送状態という生理形質を取得・解析している。
abstractlack of effective methodologies for comprehensive studies
Timely crop stress detection is essential for safeguarding yields and promoting sustainable agriculture. Traditional vegetation indices (e.g., NDVI, EVI) are widely used but remain static, crop-agnostic, and often insensitive to early stress signals. This study proposed RL-VI, a reinforcement learning-based framework that dynamically formulates vegetation indices optimized for rice stress detection. Unlike existing methods, RL-VI integrates Sentinel-2 multispectral imagery with smartphone-captured RGB data, creating the first cross-platform environment where vegetation indices are learned rather than predefined. The reinforcement learning agent adaptively selects stress-sensitive spectral band combinations guided by classification rewards. Experiments on real-world rice fields in Tamil Nadu, India, and benchmark datasets (Indian Pines, wheat salt stress) show that RL-VI achieves an overall accuracy of 89.4% and F1-score of 0.88, outperforming static and machine-learned indices by up to 12%. Importantly, RL-VI enables early stress detection up to 10 14 days before visible symptoms, providing actionable lead time for intervention. The proposed framework is computationally lightweight and scalable to UAV or edge devices, offering a farmer-ready tool for precision agriculture, bridging field-level mobile sensing with satellite monitoring for low-cost, real-time crop health management. Statistical validation using ANOVA (F = 88.24, p < 0.001) and pairwise t-tests (p < 0.001) confirmed RL-VI's superiority, while SHAP analyses emphasized the physiological significance of red-edge and SWIR bands in stress discrimination.
Why it matches plant phenotyping methods植物ストレス状態を推定する動的植生指数と強化学習フレームワークを開発し、実圃場・ベンチマークデータで性能検証しているため、フェノタイピング手法が中心である。
abstractThis study proposed RL-VI, a reinforcement learning-based framework that dynamically formulates vegetation indices optimized for rice stress detection.
Reproduction assets foundThe paper publicly releases its authors' field-captured mobile RGB rice canopy dataset on Kaggle and its full RL-VI analysis code (RL formulation, preprocessing, VI computation, training, evaluation) on GitHub. Sentinel-2 imagery and benchmark datasets are third-party public sources, not paper-specific deposits.Dataset · publicThe Mobile RGB dataset, consisting of field-captured rice canopy images collected by the authors at Polur, Tamil Nadu, India, is publicly available on Kaggle under a CC BY-NC 4.0 license (DOI: [https://doi.org/10.34740/kaggle/dsv/14105754](https:/doi.org/10.34740/kaggle/dsv/14105754)).Open asset ↗Kaggle · 10.34740/kaggle/dsv/14105754html-lines:616-683Code · publicAll custom code developed for this work including the RL-VI (Reinforcement Learning–based Vegetation Index) formulation algorithm, image preprocessing scripts, vegetation index computation modules, model training pipelines, and evaluation routines is openly accessible in a public GitHub repository. The code is available without restriction for non-commercial research use and fully available at Github Repository (https://github.com/Poornisrm/Vegetation-Index.git).Open asset ↗GitHub · Poornisrm/Vegetation-Indexhtml-lines:684-711Plant phenotyping relevance match · UnverifiedEurope PMC · checked 5 Sept 2026
The integration of Large Language Models (LLMs) with Vision-Language Models (VLMs) holds transformative potential for plant stress phenotyping, enhancing high-throughput crop monitoring, trait identification, and decision support. Traditional phenotyping methods, often reliant on manual assessments and task-specific Machine Learning (ML) models, face persistent limitations in scalability, adaptability, and contextual interpretation, especially under complex and overlapping stress conditions. VLMs address these challenges by combining deep visual recognition with contextual reasoning, enabling real-time analysis of multimodal inputs such as high-resolution imagery, agronomic text data, and environmental sensor readings. Complementarily, LLMs contribute to text mining, semantic annotation of trait descriptors, and the integration of external knowledge via Retrieval-Augmented Generation (RAG), thereby enhancing the interpretability and adaptability of phenotyping workflows. This review critically evaluates the emerging role of integrating LLMs with VLMs in plant stress phenotyping, highlighting their applications in visual trait recognition, knowledge extraction, and autonomous decision-making. We synthesize current advances and identify key challenges, including data quality, domain-specific generalization, model transparency, and equitable access to AI technologies. As one of the first comprehensive reviews on this topic, we propose a forward-looking framework that integrates LLMs, VLMs, and RAG systems to enable scalable, explainable, and user-centric phenotyping solutions. This interdisciplinary convergence offers a promising pathway toward sustainable and resilient AI-driven agriculture.
Why it matches plant phenotyping methods植物ストレス・形質認識のためのLLM/VLM統合を批判的に評価するレビューであり、植物フェノタイピング手法が中心です。
abstractThis review critically evaluates the emerging role of integrating LLMs with VLMs in plant stress phenotyping
AgriConnect+ is an AI-based decision support system designed to assist farmers in managing crop diseases, price uncertainty, and soil-driven crop selection. The framework integrates deep learning and machine learning techniques to provide three key functionalities: crop disease detection using CNN-based segmentation, crop price prediction through ensemble learning on historical market data, and soil-based crop recommendation using Top-K ranking models. Experimental results show effective disease localization, accurate price forecasting, and reliable crop recommendations. The integrated architecture enables scalable deployment and supports data-driven, sustainable agricultural decision-making.
Why it matches plant phenotyping methodsCNNベースのセグメンテーションによる作物病害の局在化を、意思決定支援システムの主要機能として評価しており、植物の病害状態を画像から推定する手法が実質的に含まれる。
abstractcrop disease detection using CNN-based segmentation
Effective pest and disease detection plays a crucial role in minimizing crop losses and improving decision-making in precision agriculture. Among the most destructive pests affecting maize crops globally is the Fall Army Worm (FAW), known for its rapid spread and high impact on yield. Existing detection practices often rely on manual scouting, which can be inefficient, labour intensive and prone to human error. This study proposes a novel deep learning based framework for the automatic classification of FAW infested and healthy maize crops by integrating RGB and thermal image modalities. The core objective is to enhance detection accuracy through multimodal image fusion. A hybrid DNN-ViT model is introduced, combining two complimentary pipelines: (i) feature-level fusion, where CNN extracted features from RGB and thermal images are fused and classified using a Deep Neural Network (DNN) and (ii) image-level fusion, where a 6 channel RGB-thermal image is directly processed using a modified Vision Transformer (ViT). Experimental results demonstrate that the fused model achieved superior performance with an accuracy of 0.98, precision, recall and F1-score of 0.98 and AUC-ROC of 0.98 on the test set, outperforming models trained on RGB-only, thermal-only and unfused data. The ablation study confirms the effectiveness of multimodal fusion, with the no-fusion model showing significantly lower performance (accuracy-0.60 and AUC-ROC-0.67). This work highlights the benefits of integrating complementary data sources for robust crop health monitoring. Future research will explore enhanced fusion strategies, environmental robustness and field level deployment to validate the model's practical applicability.
Why it matches plant phenotyping methodsRGB・熱画像融合によるFAW被害・健全状態の画像判定モデルを開発し、融合方式や性能を比較検証しているため、植物の健康状態を取得する方法が中心である。
abstractThis study proposes a novel deep learning based framework for the automatic classification of FAW infested and healthy maize crops by integrating RGB and thermal image modalities.
Reproduction assets foundThe paper's paired RGB/thermal maize FAW image dataset is publicly deposited on Figshare (part of a peer-reviewed data publication), and the authors' custom Python analysis code is released as a public supplementary file (Supplementary Code.zip) with explicit availability language. The Figshare URL matches an allowed, Dataset · publicThe dataset has been made publicly available in the Figshare Data repository as a part of a peer reviewed data publication54. Detailed information on data acquisition, sensor specifications, environmental conditions and annotation protocols is provided in the associated data article. The dataset can be accessed at: https://figshare.com/s/677d2384ba6e02db9230 (10.6084/m9.figshare.28388018).Open asset ↗Figshare · 10.6084/m9.figshare.28388018html-lines:324-345Code · publicThe custom python code developed for this study is available as supplementary file (“Supplementary Code.zip”) and includes all scripts necessary to reproduce the multimodal feature fusion, image-level fusion and ablation experiments described in the manuscript. The dataset used is publicly available on Figshare. All dependencies are listed within the code file. Readers can execute the python script to reproduce the reported results.Open asset ↗html-lines:324-345Plant phenotyping relevance match · UnverifiedCrossref · checked 5 Sept 2026
Water availability critically affects basil (Ocimum basilicum L.) growth and physiological performance, making the early and precise monitoring of water-deficit responses essential for precision irrigation. However, conventional visual or biochemical methods are destructive and unsuitable for real-time assessment. This study presents a multimodal optical biosensing and 3D convolutional neural network (3D-CNN) fusion framework for phenotyping physiological responses of basil under water-deficit stress. RGB, depth, and chlorophyll fluorescence (CF) imaging were integrated to capture complementary morphological and photosynthetic information. Through the fusion of 130 optical parameter layers, the 3D-CNN model learned spatial and temporal–spectral features associated with resistance and recovery dynamics, achieving 96.9% classification accuracy—outperforming both 2D-CNN and traditional machine-learning classifiers. Feature-space visualization using t-SNE confirmed that the learned latent representations reflected biologically meaningful stress–recovery trajectories rather than superficial visual differences. This multimodal fusion framework provides a scalable and interpretable approach for the real-time, non-destructive monitoring of crop water stress, establishing a foundation for adaptive irrigation control and intelligent environmental management in precision agriculture.
Why it matches plant phenotyping methodsバジルの水ストレス応答を、RGB・深度・クロロフィル蛍光画像と3D-CNNで非破壊推定するフェノタイピング手法が研究の中心である。
abstractThis study presents a multimodal optical biosensing and 3D convolutional neural network (3D-CNN) fusion framework for phenotyping physiological responses of basil under water-deficit stress.
Forest biomass quantification (Mg ha⁻¹) is essential for ecosystem monitoring, especially in areas under anthropogenic pressure, such as Atlantic Forest fragments. This study aimed to compare remote sensors in biomass mapping and population stock estimation of an Atlantic Forest fragment. Ten 0.1 ha plots were randomly distributed within a 17 ha fragment. Data from Sentinel-1 (S1), Sentinel-2 (S2), digital aerial photogrammetry (DAP), and their fusion were evaluated for the construction of predictive models. Linear models with two predictors were fitted: one for each sensor and another using data fusion, selecting the best predictors among all. The models were applied to estimate stand biomass using a regression estimator. The data fusion model showed the best predictive performance (RMSE = 41%), while the DAP-based model had the highest error (RMSE = 64%). However, the most accurate population estimate was obtained with the S2-based model (SE = 21 Mg ha-1), with a relative efficiency 7% higher compared to the traditional inventory (SE = 22 Mg ha-1). Estimates based on DAP, S1, and fusion were less accurate than those from the field inventory. The selected metrics, such as vegetation indices (S2) and textural metrics (S1), reflected the sensors' sensitivity to canopy structure and foliage abundance. DAP showed limitations, possibly due to its low canopy penetration. It is concluded that although the data fusion between DAP and S2 produced the best model for biomass mapping, S2 alone proved more advantageous for population estimates in forest fragments with limited sampling.
Why it matches plant phenotyping methods複数のリモートセンシング手法とデータ融合を比較し、森林キャノピー由来のバイオマス推定モデルを構築・性能評価している。植物群落の明示的な形質推定が中心であり、単なる生物学的実験の routine 測定ではない。
abstractThis study aimed to compare remote sensors in biomass mapping and population stock estimation of an Atlantic Forest fragment.
TomatoGreenhouseMultimodalRGB / grayscaleMultispectral / hyperspectralFruitClassificationSegmentationGrowth / development / phenology
Computer vision and multispectral imaging have increasingly become essential tools in modern precision agriculture. Accurate ripeness assessment is critical for yield optimization, reducing post-harvest losses, and enabling automated harvesting systems. However, traditional RGB-based approaches struggle to differentiate subtle maturity changes, and existing solutions often fail under varying lighting, occlusion, or cultivar-specific conditions. To address these challenges, this study focuses on the integration of complementary spectral cues for reliable tomato ripeness evaluation. The work utilizes a curated RGB-NIR tomato dataset comprising 224 hyperspectral samples, processed into aligned multimodal image pairs with balanced ripeness categories.The proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues. The novelty of the methodology lies in the dynamic cross-attention mechanism, which learns inter-modal dependencies between RGB and NIR signals for enhanced ripeness interpretation. Performance metrics including accuracy, precision, recall, F1-score, mIoU, and AUC were used to comprehensively evaluate the system. Experimental results demonstrate that TomatoRipen-MMT significantly outperforms all baseline RGB-only, NIR-only, and fusion methods, achieving 94.8% classification accuracy and 82.6% mIoU. These findings establish the effectiveness of multimodal Transformers for robust, high-precision fruit maturity assessment in controlled and greenhouse environments.
Why it matches plant phenotyping methodsトマト果実の成熟度という植物器官の状態を、RGB・NIR画像融合とTransformerで推定する手法を開発・評価しており、フェノタイピング手法が中心です。
abstractThe proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues.
Reproduction assets foundThe paper's phenotyping analysis is built on a publicly available USDA/NAL hyperspectral tomato dataset, explicitly linked in the Data Availability statement with an exact URL match. No author code or model checkpoints are disclosed.Dataset · publicThe dataset analyzed in this study is publicly available at the https://agdatacommons.nal.usda.gov/articles/dataset/Data_from_b_Hyperspectral_Imaging_Analysis_for_Early_Detection_of_Tomato_Bacterial_Leaf_Spot_Disease_b_/26046328.Open asset ↗26046328html-lines:1038-1053Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 6 Sept 2026
Timely detection of crop diseases in large, heterogeneous agricultural fields is difficult, as aerial imagery is often corrupted by illumination, weather, and crop-stage variations. This paper introduces AgroVisionNet, an AI-powered drone and computer vision approach that synthesises high-resolution drone imagery with in-field IoT/environmental sensor data to enhance early disease detection. The core of the proposed model is a hybrid CNN-Transformer backbone to extract spatial and contextual data from drone images, and an adaptive fusion layer to fuse time-aligned sensor readings and to make a decision, using visual and environmental evidence. Particularly, a multimodal drone–sensor dataset is collected across multiple crops and field conditions. Beyond widely used deep models for plant/crop disease identification, such as VGG16, ResNet50, Inception V3, and DenseNet121, experiments are conducted using the same training and evaluation framework. It is shown that AgroVisionNet achieves higher classification accuracy and F1-score, while inference remains feasible on an NVIDIA Jetson Nano using TensorFlow Lite. Moreover, by generating Grad-CAM plots, the study demonstrates that the proposed approach identifies disease-affected areas and, in this sense, provides interpretable information required by agronomists. These outcomes suggest that AI-based crop health tracking can be robust and field-ready by integrating drone imagery, sensor fusion, and edge computing.
Why it matches plant phenotyping methods植物の病害状態をドローン画像とセンサーから推定する手法を開発し、データセット収集、比較評価、エッジ実装、可視化まで行っており、病害フェノタイピング手法が中心である。
abstractThis paper introduces AgroVisionNet, an AI-powered drone and computer vision approach that synthesises high-resolution drone imagery with in-field IoT/environmental sensor data to enhance early disease detection.
Crop diseases pose a critical threat to global food security. Traditional diagnostic methods are inefficient and fail to meet the demands of modern precision agriculture. In recent years, artificial intelligence (AI) technologies centered on deep learning have revolutionized the rapid and precise identification of crop diseases. This paper systematically outlines key AI techniques for crop disease recognition, including computer vision-based image recognition, multimodal data fusion, and edge computing for field deployment. By analyzing representative domestic and international application cases, this paper highlights the significant advantages of this technology in terms of accuracy and efficiency. Simultaneously, it delves into current technical bottlenecks and deployment barriers, such as the few-shot learning problem, environmental interference, and low farmer trust. The paper concludes by outlining future directions, including self-supervised learning, digital twins, and industry integration, to advance the deep application and implementation of AI technology in smart agriculture.
Why it matches plant phenotyping methods作物病害を画像認識・深層学習で推定する技術を中心に、手法、精度、展開課題、将来方向を体系的にレビューしており、植物の病害状態を対象とするフェノタイピング方法レビューに該当する。
abstractThis paper systematically outlines key AI techniques for crop disease recognition, including computer vision-based image recognition, multimodal data fusion, and edge computing for field deployment.
Field / plotMultimodalLeafClassificationStress / disease detectionDisease symptoms / severity
Plant leaf diseases significantly affect crop yield and quality, posing a major challenge to global agriculture. Conventional detection methods based on manual inspection or image-only deep learning often overlook critical environmental factors influencing disease development. This paper proposes an Edge-Aware Multi-Modal Attention Network (EMAN) that integrates IoT sensor data with high-resolution leaf images to detect and assess disease severity accurately in real time. EMAN fuses visual leaf features with environmental measurements, including temperature, humidity, soil moisture, and light intensity. Dynamic Dilated Residual Blocks (DDRB) extract multi-scale image features to identify small lesions, discolorations, and subtle disease indicators. A dual attention mechanism emphasizes disease-prone regions (spatial attention) and relevant environmental factors (sensor attention), enhancing predictive performance. The architecture supports edge-cloud hybrid inference, enabling lightweight edge models to provide instant alerts to farmers, while cloud analytics refine predictions and monitor temporal disease trends. Experiments on combined Plant Village and on-field datasets demonstrate that EMAN outperforms conventional CNN-based methods in both disease classification and severity estimation, achieving robust multi-modal fusion and real-time edge inference. The proposed system provides a scalable solution for precision agriculture, promoting proactive crop management, minimizing economic losses, and supporting sustainable farming practices.
Why it matches plant phenotyping methods葉画像とIoTセンサーを統合し、植物病害の検出・重症度推定手法を開発・評価しており、植物状態の取得が研究の中心である。
abstractThis paper proposes an Edge-Aware Multi-Modal Attention Network (EMAN) that integrates IoT sensor data with high-resolution leaf images to detect and assess disease severity accurately in real time.
Detecting plant diseases is essential to contemporary agriculture since it allows for early intervention to increase crop output and reduce financial losses. Recent developments in deep learning (DL) and machine learning (ML) have shown great promise for automating the detection of diseases using sensor and visual data. Convolutional neural networks (CNNs), ResNet, DenseNet, U-Net, Mask R-CNN, and YOLO are examples of state-of-the-art DL architectures that are frequently used for extracting and learning hierarchical features from leaf pictures, hyperspectral images, and other multimodal plant datasets. This paper reviews developments in the field from 2015 to 2022. Prominent datasets like PlantVillage, Agri-Vision, and PlantDoc have made it easier to compare and test different models. Using a hybrid convolutional backbone based on EfficientNet-B7 to strike a compromise between high representational capacity and computational economy, we offer a repeatable framework that combines many cutting-edge methods for better illness detection and classification. Multiscale dilated convolutions are used to capture multiscale disease patterns, enhancing feature representations without adding more computational overhead, and adaptive segmentation mechanisms allow precise lesion localization for region-level analysis that supports severity assessment and classification. The framework emphasizes repeatability and practical application for research and industrial usage, and it is accompanied by comprehensive methodological explanations, algorithm pseudocode, and illustrated diagrams and flowcharts. In addition, the paper addresses common issues in plant disease detection, such as the lack of labeled datasets, inter-domain variability, class imbalance, and hardware constraints for real-time deployment. It also discusses mitigation strategies, such as transfer learning from pre-trained models, generative adversarial network (GAN)-based data augmentation, and model compression techniques for optimal edge deployment. In order to direct future research and real-world application, evaluation measures, performance analyses, and robustness concerns are also included. Overall, this survey and suggested methodology offer a thorough overview of current developments and solutions in ML- and DL-based plant disease detection. They show how integrating hybrid convolutional architectures, multiscale dilations, and adaptive segmentation can improve detection accuracy while addressing practical limitations, providing a scalable approach for precision agriculture and assisting researchers, agronomists, and practitioners in creating dependable, effective, and repeatable plant disease detection systems.
Why it matches plant phenotyping methods植物病害の画像認識・病斑局在化・重症度評価を扱うレビューであり、植物の病害状態を画像から抽出する方法論が中心です。
abstractThis paper reviews developments in the field from 2015 to 2022.
Plant phenotyping relevance match · UnverifiedOpenAlex · Europe PMC · checked 6 Sept 2026
Accurate measurement of key phenotypic traits, including the horizontal and vertical diameters, the weights of both fruit and pit, is essential for the selection of elite litchi cultivars and the advancement of breeding research. Manual measurement, however, is laborious, inefficient, and subjective, highlighting the urgent need for automated and precise phenotyping tools. Unlike apples, mangoes, and grapes, litchi combines a spiny, highly variable pericarp (heterogeneous areoles/tubercles across cultivars) with diverse seed morphology (including irregular, wrinkled aborted seeds), thereby increasing the difficulty of semantic segmentation and biasing diameters and weight estimation. This study presents LitchiPhenoNet, a multimodal learning framework for litchi phenotypic analysis that employs a dual-branch architecture integrating RGB (color/texture) and depth (spatial/structural) information. Experiments were conducted on an RGB-D dataset comprising 1,198 image pairs (1280×720) across 10 cultivars, using a stratified train/test split of 958/240 pairs by cultivar. To address inherent semantic and scale inconsistencies between modalities, the framework incorporates the RD-Fusion module for precise cross-modal feature extraction, improving robustness under complex and variable pericarp surfaces. Comparative experiments show that LitchiPhenoNet consistently outperforms leading YOLO-based models, achieving millimeter-level diameter estimation with coefficients of determination approaching 0.98 and mean errors within 2 mm. For weight estimation, gram-level precision is attained across whole fruit, pit, and pulp, with coefficients of determination up to 0.98 and mean errors comparable to repeated manual measurements. By handling fine-scale surface relief and cross-cultivar variability, the framework is readily extensible to other textured fruits and scalable for high-throughput phenotyping in breeding programs. Collectively, these results demonstrate that LitchiPhenoNet provides an efficient, reliable, and accurate solution for quantifying litchi phenotypic traits, substantially advancing the objectivity and efficiency of phenotypic analysis and breeding selection.
Why it matches plant phenotyping methodsRGB-D画像を用いてライチ果実・種子・果肉の径と重量を自動推定する専用フレームワークを開発し、複数品種・比較実験で性能検証しているため、植物表現型取得法が中心である。
abstractThis study presents LitchiPhenoNet, a multimodal learning framework for litchi phenotypic analysis that employs a dual-branch architecture integrating RGB (color/texture) and depth (spatial/structural) information.
Vegetation vertical structure refers to the 3D distribution of vegetation aboveground biomass. Vegetation vertical structure of tropical forests influences other ecological and environmental variables that are essential for the functioning of the ecosystems. Integrating over 5.9 million Globel Ecosystem Dynamics Investigation (GEDI) LiDAR (Light Detection and Ranging) footprints, multispectral, and synthetic aperture radar (SAR) imagery, we built five national maps at 25 m resolution of five forest structural metrics for Colombia, South America, for the year 2020. We mapped canopy height, the height of half the cumulative returned energy from GEDI (RH50), total canopy cover, foliage height diversity, and total plant area index. The resulting maps tended to have the highest errors in the Amazon and Andean regions. Total cover had the highest relative error. Interrelationship curves between forest structural metrics of GEDI footprints are maintained across mapped metrics, indicating that the predictive models preserve structural relationships observed in GEDI data. Due to the medium-high spatial resolution and national coverage of the forest structural maps presented in this work, these maps will be useful for evaluating and mapping other ecological variables and conservation priorities in Colombia.
Why it matches plant phenotyping methodsGEDI LiDAR・マルチスペクトル・SARを統合し、森林キャノピー高、被覆率、葉群高多様性、植物面積指数などの植物構造形質を全国規模で推定・検証することが中心であり、単なる生態学的応用ではない。
abstractIntegrating over 5.9 million Globel Ecosystem Dynamics Investigation (GEDI) LiDAR (Light Detection and Ranging) footprints, multispectral, and synthetic aperture radar (SAR) imagery, we built five national maps at 25 m resolution of five forest structural metrics for Colombia, South America, for the year 2020.
Reproduction assets foundThe paper's resulting forest vertical structure maps (CH, COVER, FHD, PAI, RH50 for Colombia, 2020) are publicly available on Zenodo and via Google Earth Engine assets, and the authors' analysis code is publicly available on GitHub. These are paper-specific, public, actionable assets.Code · publicCode availability
The code is publicly accessible on Github76: https://github.com/CamiloFaguaUNAL/Forest_Structure_Colombia.Open asset ↗GitHubhtml-lines:731-755Plant phenotyping relevance match · UnverifiedCrossref · OpenAlex · checked 15 Sept 2026
Field / plotMultimodalRootWhole plant / canopy / plot / fieldCountingObject detectionStress / disease detectionGrowth / time-series analysisGrowth / development / phenologyRoot system architecture
• Presents a full-process review of image-based high-throughput plant phenotyping (HTPP). • Covers recent advances in platforms, sensors, deep learning, and field-level applications. • Highlights emerging methods like Promptable models, Digital Twins, and weak supervision. • Discusses deployment challenges including data scarcity and model generalization. • Proposes future directions: multimodal fusion, uncertainty modeling, and lightweight design. With the rapid global population growth and increasing challenges in sustainable agriculture, high-throughput plant phenotyping (HTPP) has become a vital tool for advancing crop breeding and precision agriculture. This review provides a comprehensive overview of recent technological trends in image-based HTPP, focusing on the integration of advanced sensors, automated phenotyping platforms, and deep learning techniques. We summarize the evolution of imaging modalities, including 2D, 2.5D, and 3D sensors, and their respective applications in phenotype acquisition. We then examine the progress of deep learning-based models in core phenotyping tasks such as stress and disease detection, growth monitoring, organ counting, root system analysis, and postharvest quality assessment. Special attention is given to the emergence of Transformer architectures, multimodal fusion strategies, weakly supervised learning, and prompt-based foundation models. Despite significant advancements, current HTPP systems still face several challenges, including high costs, limited generalization in open-field conditions, and the need for large-scale annotated datasets. To address these, we discuss potential solutions such as transfer learning, synthetic data generation via digital twins, lightweight deployment for edge devices, and uncertainty estimation for model interpretability. By highlighting key developments and open problems, this review aims to guide future research toward scalable, robust, and intelligent plant phenotyping systems that can operate reliably in real-world agricultural environments.
Why it matches plant phenotyping methods画像ベース高スループット植物フェノタイピングのセンサー、プラットフォーム、画像解析技術を包括的にレビューしており、方法論が中心です。
abstractPresents a full-process review of image-based high-throughput plant phenotyping (HTPP).
Accurate monitoring of sunflower heads is critical for yield prediction, yet traditional methods are labor-intensive. This study proposed a novel framework integrating UAV remote sensing, deep learning, and point cloud analysis to address this challenge. The proposed method used a Dual-Branch YOLOv10n model, leveraging multi-modal data for precise detection of sunflower heads at various growth stages. Feature indices were designed and a two-step clustering technique was applied to extract sunflower head point clouds, from which geometric parameters such as diameter and volume are computed. The detection model achieved high accuracy (precision: 0.9, recall: 0.894, mAP@50: 0.932) across growth stages. A strong correlation (R² = 0.80) was found between diameter measurements from point cloud and ground-truth data, while volume showed good alignment with biomass (R² = 0.61). This method offers an innovative, efficient solution for field-scale crop monitoring and yield estimation, advancing agricultural practices.
Why it matches plant phenotyping methodsUAV画像・深層学習・点群解析を統合し、ヒマワリ頭部の検出から直径・体積という植物器官形質を抽出・検証する方法が研究の中心であるため。
abstractThis study proposed a novel framework integrating UAV remote sensing, deep learning, and point cloud analysis to address this challenge.
Precise detection of rapeseed and the growth of its canopy area are crucial phenotypic indicators of its growth status. Achieving accurate identification of the rapeseed target and its growth region provides significant data support for phenotypic analysis and breeding research. However, in natural field environments, rapeseed detection remains a substantial challenge due to the limited feature representation capabilities of RGB-only modalities. To address this challenge, this study proposes a dual-modal instance segmentation network, MSNet, based on YOLOv11n-seg, integrating both RGB and Near-Infrared (NIR) modalities. The main improvements of this network include three different fusion location strategies (frontend fusion, mid-stage fusion, and backend fusion) and the newly introduced Hierarchical Attention Fusion Block (HAFB) for multimodal feature fusion. Comparative experiments on fusion locations indicate that the mid-stage fusion strategy achieves the best balance between detection accuracy and parameter efficiency. Compared to the baseline network, the mAP50:95 improvement can reach up to 3.5 %. After introducing the HAFB module, the MSNet-H-HAFB model demonstrates a 6.5 % increase in mAP50:95 relative to the baseline network, with less than a 38 % increase in parameter count. It is noteworthy that the mid-stage fusion consistently delivered the best detection performance in all experiments, providing clear design guidance for selecting fusion locations in future multimodal networks. In addition, comparisons with various RGB-only instance segmentation models show that all the proposed MSNet-HAFB fusion models significantly outperform single-modal models in rapeseed count detection tasks, confirming the potential advantages of multispectral fusion strategies in agricultural target recognition. Finally, the MSNet was applied in an agricultural case study, including vegetation index level analysis and frost damage classification. The results show that ZN6–2836 and ZS11 were predicted as potential superior varieties, and the EVI2 vegetation index achieved the best performance in rapeseed frost damage classification.
Why it matches plant phenotyping methodsマルチスペクトル画像からナタネの個体・キャノピー領域を抽出するセグメンテーション手法を開発し、性能比較と霜害分類への応用を行っており、植物表現型取得が中心である。
abstractPrecise detection of rapeseed and the growth of its canopy area are crucial phenotypic indicators of its growth status.
MultimodalLiDAR / point cloudAnnotation / quality controlClassificationObject detectionCalibration / preprocessingSegmentationTracking
Plant phenomics, the comprehensive study of plant phenotypes, has gained prominence as a vital tool for understanding the intricate relationships between genotypes and the environment. Image-based plant phenomics has progressed rapidly, and three-dimensional (3D) phenotyping is a valuable extension of traditional 2D phenomics. However, the increased data dimensionality poses challenges to feature extraction and phenotyping. In recent decades, deep learning has led to remarkable progress in revolutionizing 3D phenotyping. Therefore, this review highlights the importance of using deep learning in 3D plant phenomics. It systematically overviews the capabilities of deep learning for 3D computer vision, covering 3D representation, classification, detection and tracking, semantic segmentation, instance segmentation, and generation. Additionally, deep learning techniques for 3D point preprocessing (e.g., annotation, downsampling, and dataset organization) and various plant phenotyping tasks are discussed. Finally, the challenges and perspectives associated with deep learning in 3D plant phenomics are summarized, including (1) benchmark dataset construction by using synthetic datasets and methods such as generative artificial intelligence and unsupervised or weakly supervised learning; (2) accurate and efficient 3D point cloud analysis by leveraging multitask learning, lightweight models, and self-supervised learning; and (3) deep learning for 3D plant phenomics by exploring interpretability, extensibility, and multimodal data utilization. The exploration of deep learning in 3D plant phenomics is poised to spur breakthroughs in a new dimension of plant science.
Why it matches plant phenotyping methods3D植物フェノミクスにおける深層学習手法を体系的にレビューしており、植物形質の抽出・推定手法が中心である。
abstractTherefore, this review highlights the importance of using deep learning in 3D plant phenomics.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Dataset · publicThe dataset can be downloaded from https://github.com/Jinlab-AiPhenomics/Mazie3D.Open asset ↗Jinlab-AiPhenomics/Mazie3Dhtml-lines:332-336Plant phenotyping relevance match · UnverifiedEurope PMC · checked 15 Sept 2026
Published1 Dec 2025Computers and Electronics in Agriculture.
Efficient and high-throughput prediction of crop yield and accurate assessment of varietal drought tolerance are essential for modern precision breeding and agricultural resource management and optimization. In this study, we propose a Stages-based Multimodal Data Fusion Model (S-MDFM) by integrating low-cost, high-throughput UAV-based multimodal imagery and multivariate data extracted from images. A staged error metrics (SEMs) propagation mechanism is constructed to capture the dynamic characteristics across growth stages and their dependencies with final yield, thereby improving the accuracy of cross-stage yield prediction (R2 = 0.8536, rRMSE = 16.12 %), validated prediction accuracy using data from the subsequent year in the same cropping area (R2 = 0.8370, rRMSE = 16.63 %). Based on wheat yield under varying water treatment gradients in drought stress experiments, the Drought Stress Tolerance Index (DSTI) and Drought Stress Susceptibility Index (DSSI) were employed to construct a Drought Resistance Index Differential (DRID) evaluation system. This dual-index approach enables multi-level screening of wheat varieties for drought resistance and quantitatively captures the synergistic relationship between yield performance and water adaptability and is capable of performing multi-level screening and quantifying the synergistic relationship between yield performance and water adaptability in wheat varieties, with an identification rate of drought-tolerant and high-yielding cultivars reaching 83.3 %. A multimodal fusion strategy based on discontinuous time-phase observations provides technical support for yield assessment and precise identification of drought-resistant and high-yielding varieties of wheat at different fertility stages, and the method provides a scalable multimodal data integration framework for varietal selection and breeding under drought-stressed field conditions, which is of great practical value for precision agriculture breeding applications.
Why it matches plant phenotyping methodsUAVマルチモーダル画像から画像特徴を抽出し、段階的データ融合モデルでコムギ収量を予測し、乾燥耐性品種を定量スクリーニングする手法が研究の中心であり、翌年データによる検証も行っている。
abstractwe propose a Stages-based Multimodal Data Fusion Model (S-MDFM) by integrating low-cost, high-throughput UAV-based multimodal imagery and multivariate data extracted from images.
Field / plotMultimodalClassificationStress / disease detectionDisease symptoms / severity
In phytopathological diagnostics, traditional unimodal, vision-based methodologies often encounter limitations in performance due to the ambiguous manifestation of disease symptoms and the complexity of background environments. In particular, the robustness and generalizability of such methods are further limited under field conditions characterized by variable illumination, occlusion, and background clutter. Conversely, textual modality exhibits robust semantic encoding capabilities, enabling precise characterization of disease phenotypes and enhancing the discriminative capability of the model. However, textual modality remains underexploited and insufficiently integrated within current research paradigms. In this study, we propose a multimodal diagnostic framework that leverages large multimodal models to automatically generate structured textual descriptions from crop imagery, thereby reducing reliance on manual annotations. To further enhance cross-modal integration, a Projected Visual–Textual Discriminant (PVD) module is introduced. Empirical results indicate that by incorporating generated textual data, most multimodal architectures consistently outperform their unimodal equivalents across diverse visual backbone networks. Notably, the combination of CogAgent with CLIP (ViT-L/14) and the PVD module achieves an F1 score of 70.76%, while LLaVA combined with ResNet50+LSTM achieves 66.38%. The proposed methodology offers a practical and scalable solution for in-field crop disease diagnosis by better aligning visual and textual representations, enhancing classification performance, and obviating the need for manually curated textual annotations by automatically generating descriptions using the Automated Image Description Generation module.
Why it matches plant phenotyping methods作物画像から病徴を推定するマルチモーダル診断フレームワークを開発し、分類性能を比較評価しており、植物病害状態の表現型取得・推定が中心である。
abstractwe propose a multimodal diagnostic framework that leverages large multimodal models to automatically generate structured textual descriptions from crop imagery
Accurate crop monitoring is essential for agricultural planning and food security. This study developed a coupling framework of unmanned aerial vehicle (UAV) multimodal data and crop models based on a sequential data assimilation method, offering technical support for crop growth simulation and precision management under drip irrigation modes in the Hexi Corridor of Northwest China. Multispectral and thermal infrared image data of spring maize at different growth stages were acquired via UAVs. The UAV-derived leaf area index (LAI) and soil moisture (SM) were assimilated into the WOFOST model using the ensemble Kalman filter (EnKF). Three assimilation schemes including (a) LAI, (b) SM, and (c) LAI+SM were compared to explore the effects of different mulching treatments (mulched vs. non-mulched) and irrigation gradients on assimilation performance under drip irrigation modes. Our results showed that the fusion of UAV-based multispectral and thermal infrared multimodal data enabled accurate retrieval of LAI and SM, with a maximum R² of 0.85. The three assimilation schemes exhibited significant differences, and the joint assimilation of LAI and SM outperformed the others. This may be since LAI and SM, as key indicators of crop growth and development, undergo dynamic changes throughout the growth period, and their joint assimilation fully captures the temporal variability of crops and soil. In addition, the proposed framework demonstrated marked variations in simulation accuracy across different drip irrigation modes. Overall, the performance for shallow buried drip irrigation (SBDI) was superior to that for surface drip irrigation (SDI) and film-mulched drip irrigation (FDI). This may be attributed to the direct influence on soil evaporation and evapotranspiration under the latter two modes, which in turn modifies crop growth and development processes and ultimately affects the model's simulation accuracy.
Why it matches plant phenotyping methodsUAVマルチスペクトル・熱赤外データからLAIを推定し、作物モデルへ同化する手法の開発と性能比較が研究の中心であり、植物形質取得の技術的評価を含む。
abstractThis study developed a coupling framework of unmanned aerial vehicle (UAV) multimodal data and crop models based on a sequential data assimilation method
Published1 Dec 2025Journal of food composition and analysis : an official publication of the United Nations University, International Network of Food Data Systems
The main pathogen of sweet potato black rot, Ceratocystis fimbriata, induces the production of toxic secondary metabolites, leading to significant post-harvest economic losses. Establishing early rapid detection technologies is crucial for ensuring sweet potato food safety and reducing economic losses. This study systematically monitored the spectral image (HSI), electronic nose (E-nose) response signal, and total phenolic content (TPC) reference value of sweet potato samples after artificial inoculation with the pathogen, aiming to utilize TPC as a key biochemical indicator for the early prediction of disease progression. The experiment compared single-source and multi-source data fusion methods. Results showed that the CARS-PCA-MHA-CNN model(Parameters was reduced by 96.58%) based on a feature-level fusion strategy achieved the best predictive performance (R²=0.974, RMSEP=0.041, RPD=6.14). Compared with single-source data, the prediction accuracy was improved by 7.6% and 6.4%, respectively. Furthermore, the model's generalization ability was tested on an independent test set (unenhanced). This study proposes a reliable and non-destructive method for the early prediction of postharvest diseases in root and tuber crops, which has great application potential in the field of intelligent monitoring of agricultural products.
Why it matches plant phenotyping methods病原体接種後のサツマイモの病害進行を、HSI・電子鼻・TPCデータ融合とCNNで非破壊予測する方法が研究の中心であり、植物器官の病害状態を推定する実質的なフェノタイピング手法である。
abstractThe experiment compared single-source and multi-source data fusion methods.
We constructed a computational methodology to assess health of plant-microbiome system through microbiome structure modelling combined with plant remote sensing. As a test dataset, we selected soil mycobiome and morphometry of Tilia cordata in nursery and forest sites. Our method is also applicable on forest or regional scale. Microbiome part called GiaC ( G u i lds a nd o C currences) combines taxonomic and trophic composition as well as species co-occurrence modelled with advanced graph methods. We complemented state-of-the-art approaches with novel ones for visualisations, species filtering (Flexible99) and graph transformation modelling species clusters (ClusterCollapse). Flexible99 is a method that adjusts the species abundance cut-off to each sample set and removes rare species. ClusterCollapse generalises co-occurrence networks to species clusters by edge contraction and serves as an implicit homogeneity test. To assess biomass of the seedlings we used low-cost and field-adopted morphometric and manual measurements. Top and side tree images, acquired with handheld RGB camera, were analysed using colour segmentation and pixel count based methods. Parameters, such as crown size, shape, area and pigment content, number of leaves, branch length and foliage density, allowed the seedlings to be classified into three different vitality groups. Presented multimodal approach was capable to differentiate and characterize distinct best, suboptimal or critical states of microbiome-host system, both on microbial and plant side. Our results show that more stable fungal co-occurrence patterns should be attributed to the plant set of the best growth. In contrast, more chaotic patterns can be considered non-optimal for plant-mycobiome cooperation.
Why it matches plant phenotyping methods植物の健康・活力状態を推定するマルチモーダル手法の一部として、RGB画像の色分割・画素計数から樹冠形状、葉数、枝長、葉密度などの形質を抽出しており、フェノタイピング手法の適用が実質的に含まれる。
abstractWe constructed a computational methodology to assess health of plant-microbiome system through microbiome structure modelling combined with plant remote sensing.
Turmeric and ginger are economically vital spice crops cultivated for their underground rhizomes, which are largely susceptible to conditions similar as soft spoilage, rhizome spoilage, and bacterial wilt. Traditional discovery styles calculate on visual examination orpost- harvest opinion, frequently performing in delayed treatment and significant yield loss. This exploration proposes an AI- driven frame for early rhizome complaint discovery using a multimodal approach that integrates deep literacy, hyperspectral imaging, and IoT- grounded environmental seeing. Convolutional Neural Networks (CNNs), enhanced through transfer literacy, are employed to classify rhizome health from subterranean image data, while detector emulsion ways relate soil humidity, temperature, and pH with complaint onset. The system also incorporates time- series soothsaying and natural language interfaces to deliver real- time cautions and treatment recommendations to growers. By fastening on rhizome- position analysis — an area largely overlooked in being literature — this study aims to ameliorate individual delicacy, reduce crop losses, and promote sustainable spice husbandry through intelligent, accessible technology.
Why it matches plant phenotyping methods地下部画像からショウガ・ウコンの健康状態を分類し、環境センサーと組み合わせて病害 onset を推定するAI手法が研究の中心であり、植物の病害状態を直接評価するため適格。
abstractThis exploration proposes an AI- driven frame for early rhizome complaint discovery using a multimodal approach that integrates deep literacy, hyperspectral imaging, and IoT- grounded environmental seeing.
The purpose of this study is to develop and evaluate a multi-modal fusion model that integrates image and environmental data collected from a vertical farm to accurately predict crop growth stages and growth rates. RGB images and key environmental parameters, including CO₂ concentration, temperature, relative humidity, and light intensity, were synchronously collected over a defined period at a vertical farm test site, and a multi-modal training dataset was constructed through preprocessing and feature extraction. Leaf area, greenness, shape indices, and proxy VI were extracted from the image data, while the temporal variations of the environmental sensor data were modeled using a long short-term memory (LSTM) network. These were then merged with convolutional neural network (CNN)–based image features to perform simultaneous growth stage classification and growth rate prediction.
Why it matches plant phenotyping methods画像特徴量と環境センサーデータを融合し、葉面積・緑色度・形状指標などの植物形質から生育段階と生育速度を推定する手法の開発・評価が中心である。
abstractThe purpose of this study is to develop and evaluate a multi-modal fusion model that integrates image and environmental data collected from a vertical farm to accurately predict crop growth stages and growth rates.
The accuracy of automated plant disease diagnosis is frequently limited by the use of visual symptoms alone, especially when it comes to differentiating between conditions that have a lot of visual similarities. To address this, we propose a new privacy-preserving framework that combines the strengths of multi-modal federated learning (FL) with environmental context. Our system integrates leaf images with synthetic sensor data—such as temperature, humidity, and leaf wetness duration—capturing critical cues that influence disease progression. Actually, this system core is a dualbranch convolutional neural network designed to process both image and environmental features in a way that reflects the biological characteristics of different diseases. Results demonstrates that the multi-modal approach consistently outperforms conventional image-only models across multiple disease categories, and especially true for diseases where environmental factors are very important in how they develop. We further extend the system into a federated learning setting, allowing models to benefit from distributed training while keeping sensitive agricultural data local and private. This makes the framework not only more accurate but also practical for real-world use, where data privacy is essential.
Why it matches plant phenotyping methods葉画像と環境センサーデータから植物病害状態を推定するマルチモーダル分類手法を開発し、画像のみのモデルと比較評価しており、植物フェノタイピング手法が中心である。
abstractwe propose a new privacy-preserving framework that combines the strengths of multi-modal federated learning (FL) with environmental context.
Field / plotMultimodalMultispectral / hyperspectralRaman / spectroscopyWhole plant / canopy / plot / fieldCalibration / preprocessingGrowth / development / phenology
This data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo), a native species from the arid and semi-arid regions of Southern Zacatecas, Mexico. In particular, Stenocereus spp. are important cacti in the region due to its nutritional properties, role as an economic resource, and cultural significance.It is worth emphasising that these cacti traditionally grow wild (i.e., without deliberate cultivation); accordingly, controlled cultivation is uncommon and remains understudied. With the aim of producing a formal, comprehensive analysis and compendium, the data were collected across multiple phenological stages to provide a complete representation of the plant development cycle, from vegetative growth through to fruiting. To achieve this, the collection process combined high-resolution multispectral imaging with field spectrometry in the 400-700 nm range. Standardized acquisition protocols were applied in field conditions to capture consistent reflectance data, and environmental variables such as illumination, temperature, and geographic coordinates were recorded for each session to ensure reproducibility. The dataset integrates several components: (i) multispectral images that provide spatial information on canopy and structural characteristics, (ii) field spectral signatures with detailed reflectance values for each sampled plant, and (iii) metadata describing phenological stage, acquisition date and time, environmental conditions, and equipment settings. For subsequent analysis, data was preprocessed and normalized to enable reliable comparisons between growth stages and across acquisition sessions, resulting in a clean, structured resource ready for computational analysis. In this regard, this dataset has been organized to facilitate its direct application across multiple research and development contexts. Specifically, potential applications include the training and validation of machine learning and computer vision models for automated phenological stage classification, harvest time estimation, and development of species-specific vegetation indices. Moreover, owing to its standardized design, the resource can serve as a benchmark for comparing methods, validating algorithms, and supporting reproducible workflows in precision agriculture and remote sensing. Beyond Stenocereus queretaroensis, the documented acquisition and preprocessing methodology can be replicated or adapted to generate similar multimodal datasets for other climate-resilient crops, particularly those cultivated in arid and semi-arid regions. This could enable comparative analyses across species and provide a reference for extending multimodal sensing approaches to underrepresented plants of ecological and economic importance.
Why it matches plant phenotyping methods植物の生育段階・生理・構造特性を対象に、標準化されたマルチスペクトル画像とフィールド分光データを収集・前処理した再利用可能なデータセットであり、ベンチマークやアルゴリズム検証を目的とするため、フェノタイピング手法が中心です。
abstractThis data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo)
Reproduction assets foundThe paper's own multimodal phenotyping dataset (multispectral/RGB images, spectral signatures, NDVI products, metadata, and example MATLAB scripts) is publicly deposited on Mendeley Data with explicit direct URL and DOI.Dataset · public) at ∼1750 m a.s.l., under semi-arid temperate conditions with spring temperatures ranging 20–33°C. The data were collected from the Unit Academic of Electrical Engineering Plantel Jalpa.
Data accessibility
Repository name: Multimodal_Cactaceae_Dataset_25
Data identification number: doi:10.17632/skw8tjc82f.1
Direct URL to data: https://data.mendeley.com/datasets/skw8tjc82f/1
Instructions for accessing these data: click on the direct URL to obtain the multimodal data from Mendeley Dataset Repository.
Related research article
None
1.
Value of the Data
•
These data provide a unique, non-invasive resource for studying Stenocereus spp. physiology. The integrated collection of high-resolution multOpen asset ↗Mendeley Data · doi:10.17632/skw8tjc82f.1lines:32-58Code / dataset availability confirmedEurope PMC · OpenAlex · checked 6 Sept 2026
TomatoMultimodalStereoFruitObject detectionVisualization / data managementGrowth / development / phenologyFruit / seed / panicle traitsYield / yield components
Introduction The advancement of smart agriculture has witnessed increasing applications of computer vision in crop monitoring and management. However, existing approaches remain challenged by high computational complexity, limited real-time capability, and poor multi-task coordination in tomato cultivation scenarios. Methods To address these limitations, an intelligent tomato management system is proposed based on the Ghost-based Adaptive Efficient You Only Look Once (GAE-YOLO) algorithm. The lightweight architecture of the GAE-YOLO framework is achieved through the replacement of standard convolutional layers with Ghost Convolution (GhostConv) modules, while detection accuracy is significantly improved by the integration of both AReLU activation functions and Effective Intersection over Union (E-IoU) loss optimization. The system, implemented on a Jetson TX2 embedded platform, also incorporates ZED stereo vision for 3D localization and a PyQt6-based visualization platform. Results When implemented on Jetson TX2, the system achieving 93.5% mean Average Precision at 50% intersection over union (mAP@50) at 10.2 frames per second (FPS), which can be optimized to 27 FPS by employing TensorRT acceleration and 720p resolution for scenarios demanding higher throughput. Furthermore, it establishes standardized assessment systems for tomato maturity and yield prediction, and offers integrated modules for disease diagnosis and agricultural large language model consultation. Discussion This work establishes a new paradigm for edge computing in agriculture while providing critical technical support for smart farming development.
Why it matches plant phenotyping methodsトマトの成熟度・収量予測および病害診断を含む画像・3Dビジョン基盤を開発し、エッジ環境で性能評価しているため、植物表現型取得が中心的な研究である。
abstractan intelligent tomato management system is proposed based on the Ghost-based Adaptive Efficient You Only Look Once (GAE-YOLO) algorithm
Reproduction assets foundThe paper's data availability statement explicitly states that the data and code supporting the study are publicly available on GitHub at the authors' repository (GAE-YOLO), which matches an allowed URL. This qualifies as a paper-specific public code asset for the tomato detection/phenotyping analysis.Code · publicThe data and code supporting this study are publicly available at GitHub under the following links: https://github.com/NSSCk/GAE-YOLO .Open asset ↗NSSCk/GAE-YOLOlines:756-834Plant phenotyping relevance match · UnverifiedEurope PMC · Crossref · checked 6 Sept 2026
Abstract This work presents a lightweight imaging-based AI model for rapid, point-of-care diagnosis of potato leaf diseases. Using the PlantVillage dataset comprising approximately 3,000 labeled images across three classes— Healthy , Early Blight , and Late Blight —a transfer learning approach was implemented with the MobileNetV2architecture. The dataset was split into 80% training and 20% validation sets, with preprocessing and augmentation to enhance generalization. Trained for 10 epochs using the Adam optimizer (learning rate = 0.001), the model achieved a training accuracy of 96.8% and a validation accuracy of 93.7% , with respective losses of 0.13and 0.21 . Class-wise evaluation confirmed balanced precision and recall across all categories, while external testing yielded correct disease identification with 57.2% confidence . The model demonstrates that high diagnostic accuracy can be achieved on basic hardware, making it suitable for low-resource agricultural settings. Compared to complex multimodal architectures, this MobileNetV2-based design offers fast inference, minimal computational demand, and strong generalization—establishing an efficient foundation for real-time, AI-driven plant disease diagnostics.
Why it matches plant phenotyping methodsジャガイモ葉の病害状態を画像から推定するAI診断モデルの開発・検証が研究の中心であり、植物表現型の取得方法に該当する。
abstractThis work presents a lightweight imaging-based AI model for rapid, point-of-care diagnosis of potato leaf diseases.
In the arid cultivation region of Xinjiang, China, shrinkage disease severely compromises the quality, yield, and market value of jujube. Published research has achieved high accuracy in detecting larger lesions using RGB imaging and hyperspectral imaging (HSI). However, these methods lack sensitivity in detecting early and subtle symptoms of disease. In this study, a multi-source data fusion strategy combining RGB imaging and HSI was proposed for non-destructive and high-precision detection of early-stage jujube shrinkage disease. Firstly, a total of 317 fruits of the 'Junzao' cultivar were collected during multiple stages of natural infection, covering early-stage shrinkage disease detection across different growth stages, including both green and mature red fruits. Secondly, morphological features were extracted from RGB images in multiple dimensions, while a three-stage feature selection strategy combining Principal Component Analysis (PCA), the Successive Projections Algorithm (SPA), and the Genetic Algorithm (GA) was implemented to identify four key wavelengths from HSI. Thirdly, a hybrid convolutional neural network-multilayer perceptron (CNN-MLP) architecture was constructed, with dynamic feature weighting employed to achieve effective multimodal fusion and optimize detection performance. Experimental results demonstrated that compared to the MLP and CNN models, the proposed method achieved approximately 8.0% and 5.4% improvements in accuracy and 38.6% and 32.4% improvements in F1 scores, respectively. It offers a robust and scalable solution for early disease detection and postharvest quality assessment in jujube production.
Why it matches plant phenotyping methodsRGB画像・HSIから果実の病斑形態と分光特徴を抽出し、マルチモーダル深層学習で植物病害状態を検出する手法の開発・性能評価が中心であるため。
abstracta multi-source data fusion strategy combining RGB imaging and HSI was proposed for non-destructive and high-precision detection of early-stage jujube shrinkage disease.
Reproduction assets foundThe paper's Data Availability Statement points to a public GitHub repository containing the study's dataset (RGB images and hyperspectral data of jujube fruits). No separate analysis code availability is stated, but the deposited dataset is a paper-specific, publicly actionable asset.Dataset · publicThe data from this study are publicly available. The dataset is available at https://github.com/2484733079/Early-detection-of-Jujube-Shrinkage-Disease-by-Multi-source-Data-on-Multi-task-Deep-Network.git (accessed on 13 October 2025).Open asset ↗2484733079/Early-detection-of-Jujube-Shrinkage-Disease-by-Multi-source-Data-on-Multi-task-Deep-Networklines:365-367Plant phenotyping relevance match · UnverifiedEurope PMC · checked 15 Sept 2026
Accurate assessment of cotton defoliation (DF) and boll opening (BO) is essential for optimizing yield and fiber quality during mechanized harvesting, as improper timing can reduce yield and impair fiber quality. Unmanned aerial vehicle (UAV)-based remote sensing has become an effective tool for monitoring these indicators, but most current methods rely on single-sensor data, limiting diagnostic accuracy and generalizability. To address this limitation, we propose a multi-source data fusion framework integrating RGB, multi-spectral (MS), and thermal infrared (TIR) sensors for comprehensive canopy information. The fused dataset includes vegetation indices (VIs), color indices (CIs), texture features (Tex), and canopy temperature (TC). Feature selection was performed using pearson correlation coefficients (PCCs), recursive feature elimination with cross-validation (RFECV), and the Boruta algorithm to identify key variables. Three machine learning models—partial least-squares regression (PLSR), random forest regression (RFR), and extreme gradient boosting regression (XGBR)—were developed and compared. The RFECV-selected RGB+MS+TIR features in the XGBR model achieved the highest predictive accuracy, with R² values of 0.918 for defoliation rate and 0.867 for boll opening rate, improving by 1.9 % and 4.3 %, respectively, over single-sensor models. Root mean square error (RMSE) and relative RMSE (rRMSE) were reduced by 1.11 %-1.99 % and 1.66 %-2.28 %, respectively. These findings demonstrate that multi-source UAV data fusion, combined with advanced machine learning techniques, significantly enhances the accuracy and robustness of cotton defoliation and boll opening diagnosis. This approach offers a practical solution for precision agriculture to improve harvest scheduling and defoliant management.
Why it matches plant phenotyping methodsUAVのRGB・マルチスペクトル・熱赤外データを融合し、綿花の落葉率と綿花開絮率という植物状態を推定する手法を開発・比較しており、フェノタイピング手法が研究の中心である。
abstractwe propose a multi-source data fusion framework integrating RGB, multi-spectral (MS), and thermal infrared (TIR) sensors for comprehensive canopy information.
To enhance fruit yield and quality, this study focuses on precise pre-harvest ripeness assessment and early disease detection. Addressing the limitations of conventional methods, we propose a multimodal flexible sensing and deep learning-based evaluation framework. The developed flexible optoelectronic in-situ sensing system integrates spectral (410-940 nm, 18 channels) and impedance (100 Hz-10 kHz) detection, allowing conformal attachment to mango surfaces for nondestructive monitoring throughout the growth cycle while collecting spectral, impedance, and physicochemical data. The proposed 1DCNN-ATT-BiLSTM-ATT network employs independent branches to extract local features from each modality, followed by attention mechanisms and temporal modelling for comprehensive feature fusion, achieving 97.5 % accuracy on test sets. Field experiments reveal systematic variations in soluble solid content (SSC), moisture content (MC), and optoelectronic signals during ripening. Correlation and Granger causality analyses underscore the necessity of multimodal fusion. This system supports intelligent harvesting and precision monitoring, advancing agricultural practices toward greater efficiency and sustainability while establishing a technical paradigm for precision agriculture. Future work will focus on improving environmental robustness and cross-cultivar applicability.
Why it matches plant phenotyping methodsマンゴー果実に装着する分光・インピーダンス統合センシングと深層学習による成熟度・品質状態推定を開発しており、植物状態の取得・抽出手法が研究の中心である。
abstractThe developed flexible optoelectronic in-situ sensing system integrates spectral (410-940 nm, 18 channels) and impedance (100 Hz-10 kHz) detection, allowing conformal attachment to mango surfaces for nondestructive monitoring throughout the growth cycle
Global crop losses of 20–40% continue because traditional plant assessment methods are either invasive, damaging plant tissues, or reactive, detecting stress only after visible symptoms. Recent developments have remained fragmented, focusing on single modalities, individual organs, or limited frequency ranges. This study developed a unified bioelectrical sensor system capable of non-invasive, multimodal, multiscale, and integrative assessment by integrating capabilities that existing methods address only separately. The system combines spectroscopy and tomography within a single platform, enabling simultaneous evaluation of multiple organs. Unlike approaches confined to narrow frequencies, it captures complete physiological responses across scales. Validation on strawberry (Fragaria × ananassa ‘Sweet Charlie’) demonstrated comprehensive multi-organ assessment: 98.3% accuracy for fruit categorization, 95.8% for leaf water status, and 88.2% for stem productivity. Tomographic performance reached 2.6–2.8 mm resolution for 3D root mapping and 2.8–3.0 mm for 2D postharvest fruit sorting. Correlations with reference metrics were used exclusively for validation, confirming that the extracted features reflect genuine physiological variations. Importantly, the system detects stress before visible symptoms, enabling intervention within the reversible window. By unifying spectroscopy and tomography with complete frequency coverage and multi-organ capability, this platform overcomes existing fragmentation and establishes a foundation for proactive, comprehensive plant monitoring essential for sustainable agriculture.
Why it matches plant phenotyping methods植物の生理状態を非侵襲的に取得するマルチモーダル・マルチスケール生体電気センサー基盤を開発し、果実・葉・茎・根の評価と基準指標による検証を行っており、表現型取得法が研究の中心です。
abstractThis study developed a unified bioelectrical sensor system capable of non-invasive, multimodal, multiscale, and integrative assessment
Plant diseases pose a significant threat to global food security, particularly in regions that rely heavily on crops that are vulnerable to disease, such as tomatoes. This research addresses the inefficiencies of traditional farming solutions by presenting a novel multimodal deep learning algorithm. The algorithm leverages EfficientNetB0 for image-based disease classification and utilizes Recurrent Neural Networks (RNN) to predict disease severity based on environmental data. By integrating visual and climatological inputs, our model addresses the limitations of unimodal systems, enhancing classification accuracy and interpretability. The model achieved a disease classification accuracy of 96.40% and a severity prediction accuracy of 99.20%. Additionally, the use of LIME and SHAP explainable AI techniques improves the understanding of disease severity classification outcomes. The contributions of this study align with precision agriculture practices and advance the resilience of local food systems, particularly in economies heavily dependent on tomato production. The proposed approach has the potential to mitigate the impacts of plant diseases and enhance food security by utilizing innovative technological solutions.
Why it matches plant phenotyping methodsトマト病害の画像分類と病害重症度推定を行うマルチモーダル手法が研究の中心であり、植物状態の測定・推定に直接関わるため含める。
abstractpresenting a novel multimodal deep learning algorithm
Improving light-use efficiency (LUE) is essential for boosting crop productivity, particularly in controlled-environment agriculture. Despite recent advances, most studies still rely on destructive measurements or one-dimensional data, which limits insight into the structural–physiological coordination underlying LUE. We established a multimodal phenotyping platform to dissect the phenotypic regulatory network of LUE in lettuce ( Lactuca sativa L.). Integrating hyperspectral imaging with multiview three-dimensional (3D) reconstruction, we developed a noninvasive, high-throughput system that simultaneously estimates 3D plant architecture, photosynthetic physiology—net photosynthetic rate (A) and relative chlorophyll content (SPAD)—and aboveground biomass (AGB) across 35 cultivars. A modeling pipeline combining StandardScaler (SS) normalization, genetic algorithm (GA) feature selection, and artificial neural networks (ANN) achieved robust prediction of A (R²=0.72), SPAD (R²=0.87), and AGB (R²=0.85). Spectral contribution analysis revealed distinct sensitivities: SPAD across 400–700 nm, A near 430 and 680 nm, and AGB across 500–580 nm. The 426–430 nm blue band emerged as a key region: high-efficiency cultivars showed distinctive reflectance (42.93–59.03 %), consistent with superior photosynthetic performance. Structurally, high-efficiency types exhibited “large-and-loose” canopies, with greater plant height (+64.37 %), projected area (+59.42 %), and convex-hull volume (+166.3 %), alongside reduced compactness (−23.48 %). Network analysis indicated progressively tighter coupling between spectral and structural traits from low- to high-efficiency groups, consistent with adaptive coordination for light capture and use. These results identify actionable phenotypic markers for selecting high-LUE cultivars and provide a transferable platform for phenomics-driven breeding and management in controlled-environment crops. • A multimodal framework enables non-destructive, high-throughput phenotyping in lettuce. • 66 key spectral and structural features linked to light-use efficiency were identified. • A photosynthetic trait network reveals coordination of pigments and canopy architecture. • Breeding targets for blue-light response and canopy structure optimization are proposed.
Why it matches plant phenotyping methodsレタスの構造・生理形質を推定するマルチモーダル表現型プラットフォームを開発し、非破壊・高速測定と予測性能を評価しており、表現型取得手法が研究の中心である。
abstractWe established a multimodal phenotyping platform to dissect the phenotypic regulatory network of LUE in lettuce ( Lactuca sativa L.).
Abstract Magnetic resonance imaging (MRI), long established in medical diagnostics, offers powerful, non-invasive capabilities for visualizing physiological processes in intact plants. This review focuses on the principles, recent advances, and future prospects of MRI-based lipid analysis in plant science with a particular focus on seeds. Cutting-edge, spatially resolved MRI has uncovered a remarkable compartmentation of lipid metabolism and storage. Lipid distribution patterns reflect the tissue- and cell-specific functional roles of lipids and are shaped by local metabolite gradients and other regulatory factors, including biomechanical and environmental stimuli. Recent innovations in MRI methodology now allow comprehensive, non-invasive monitoring of lipid storage and degradation dynamics in vivo. Looking ahead, the integration of MRI with deep learning and multimodal approaches heralds a transformative era for seed biology, oilseed phenotyping, and breeding.
Why it matches plant phenotyping methods植物の脂質分布・貯蔵・分解動態をMRIで非侵襲的に可視化・追跡する方法を扱うレビューであり、種子フェノタイピングへの応用も明示されているため、方法論が中心です。
abstractThis review focuses on the principles, recent advances, and future prospects of MRI-based lipid analysis in plant science with a particular focus on seeds.
Field / plotMultimodalFlowerFruitObject detectionYield / biomass estimationYield / yield components
The article presents a comprehensive system for forecasting orchard yields based on multimodal remote monitoring data. It combines convolutional neural networks for detecting flowers, ovaries, and fruits with an ensemble of linear and nonlinear models (multivariate regression, MLP, LSTM) for yield estimation. LASSO regression and SHAP analysis are used to interpret the results. The developed Python software enables full data processing, visualization, and saving of forecasts. The model achieves a determination coefficient of R2>0.85 and RMSE
Why it matches plant phenotyping methods果実園の収量予測を目的とするが、花・子房・果実をCNNで検出し、マルチモーダルデータを統合して収量を推定する取得・解析システムとPythonソフトウェアが中心であり、植物器官および収量形質の計測ワークフローに該当する。
abstractIt combines convolutional neural networks for detecting flowers, ovaries, and fruits with an ensemble of linear and nonlinear models (multivariate regression, MLP, LSTM) for yield estimation.
Securing agricultural productivity and food security from serious threats is made possible through timely and reliable disease detection. The early detection of leaf disease is revolutionized by the recent advancements in imaging technology like digital imaging and remote sensing (RS) integrated with Artificial intelligence (AI). High- resolution, close-range visual data was offered by digital imaging, and it is crucial for detecting subtle symptoms in various applications. Large-scale monitoring over fields was facilitated by the RS platforms like drones and satellites, and it may help the farmers in offering actionable insights. Modern methods for crop leaf disease detection was reviewed in this study by analysing various imaging techniques like pre-processing, feature extraction (FE), classification techniques, and deep learning (DL) advancements. The datasets included, risks, and performance metrics utilized are all discussed in this review. The field of disease surveillance using AI and multimodal imaging has shown advancements in facilitating scalable, and real-time disease surveillance, and it is also highlighted in this review. The future paths for precision agriculture are guided by conclusion of the study, as it offers insights regarding present research gaps
Why it matches plant phenotyping methods作物葉の病徴を画像から検出する手法について、画像処理・特徴抽出・分類・深層学習・データセット・性能指標を体系的に扱うレビューであり、植物表現型取得手法が中心です。
abstractModern methods for crop leaf disease detection was reviewed in this study by analysing various imaging techniques like pre-processing, feature extraction (FE), classification techniques, and deep learning (DL) advancements.
Plant diseases remain a major constraint on crop productivity, requiring timely and accurate diagnostic approaches to secure agricultural yields. While existing automated diagnosis methods primarily rely on image data and achieve notable results, their performance often declines in complex field environments with noise and interference. Multimodal learning provides a promising solution by integrating complementary cues from various data sources. However, the heterogeneity between plant phenotypes and other modalities, such as textual descriptions, poses a significant challenge for effective fusion. To address this issue, we propose PlantIF, a multimodal feature interactive fusion model for plant disease diagnosis based on graph learning. PlantIF comprises three key components: image and text feature extractors, semantic space encoders, and a multimodal feature fusion module. Specifically, we employ pre-trained image and text feature extractors to extract visual and textual features enriched with prior knowledge of plant diseases. Semantic space encoders then map these features into both shared and modality-specific spaces, enabling the capture of cross-modal and unique semantic information. To enhance context understanding, we design a multimodal feature fusion module to process and fuse different modal semantic information, and then extract the spatial dependency between plant phenotype and text semantics through the self-attention graph convolution network. We evaluate PlantIF on a multimodal plant disease dataset with 205,007 images and 410,014 texts, achieving 96.95 % accuracy, 1.49 % higher than existing models. These results demonstrate the potential of multimodal learning in plant disease diagnosis and highlight PlantIF's value in precision agriculture. Codes are available at https://github.com/GZU-SAMLab/PlantIF.
Why it matches plant phenotyping methods植物画像とテキストを統合して病害状態を診断するモデルPlantIFを開発・評価しており、植物病害という観察可能な植物状態の推定が研究の中心である。
abstractwe propose PlantIF, a multimodal feature interactive fusion model for plant disease diagnosis based on graph learning.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Code · publicCodes are available at https://github.com/GZU-SAMLab/PlantIF .Open asset ↗GZU-SAMLab/PlantIFlines:1-28Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Recent advances in hyperspectral imaging (HSI) and multimodal deep learning have opened new opportunities for crop health analysis; however, most existing models remain limited by dataset scope, lack of interpretability, and weak cross-domain generalization. To overcome these limitations, this study introduces Agri-DSSA, a novel Dual Self-Supervised Attention (DSSA) framework that simultaneously models spectral and spatial dependencies through two complementary self-attention branches. The proposed architecture enables robust and interpretable feature learning across heterogeneous data sources, facilitating the estimation of spectral proxies of chlorophyll content, plant vigor, and disease stress indicators rather than direct physiological measurements. Experiments were performed on seven publicly available benchmark datasets encompassing diverse spectral and visual domains: three hyperspectral datasets (Indian Pines with 16 classes and 10,366 labeled samples; Pavia University with 9 classes and 42,776 samples; and Kennedy Space Center with 13 classes and 5211 samples), two plant disease datasets (PlantVillage with 54,000 labeled leaf images covering 38 diseases across 14 crop species, and the New Plant Diseases dataset with over 30,000 field images captured under natural conditions), and two chlorophyll content datasets (the Global Leaf Chlorophyll Content Dataset (GLCC), derived from MERIS and OLCI satellite data between 2003–2020, and the Leaf Chlorophyll Content Dataset for Crops, which includes paired spectrophotometric and multispectral measurements collected from multiple crop species). To ensure statistical rigor and spatial independence, a block-based spatial cross-validation scheme was employed across five independent runs with fixed random seeds. Model performance was evaluated using R2, RMSE, F1-score, AUC-ROC, and AUC-PR, each reported as mean ± standard deviation with 95% confidence intervals. Results show that Agri-DSSA consistently outperforms baseline models (PLSR, RF, 3D-CNN, and HybridSN), achieving up to R2=0.86 for chlorophyll content estimation and F1-scores above 0.95 for plant disease detection. The attention distributions highlight physiologically meaningful spectral regions (550–710 nm) associated with chlorophyll absorption, confirming the interpretability of the model’s learned representations. This study serves as a methodological foundation for UAV-based and field-deployable crop monitoring systems. By unifying hyperspectral, chlorophyll, and visual disease datasets, Agri-DSSA provides an interpretable and generalizable framework for proxy-based vegetation stress estimation. Future work will extend the model to real UAV campaigns and in-field spectrophotometric validation to achieve full agronomic reliability.
Why it matches plant phenotyping methods植物のクロロフィル含量・活力・病害ストレスを画像/ハイパースペクトルから推定する新規深層学習フレームワークを開発・評価しており、植物表現型の取得・推定が中心である。
abstractthis study introduces Agri-DSSA, a novel Dual Self-Supervised Attention (DSSA) framework that simultaneously models spectral and spatial dependencies through two complementary self-attention branches.
Reproduction assets foundThe paper's Data Availability Statement explicitly deposits the authors' Agri-DSSA implementation (the computational analysis code for the phenotyping experiments) in a public GitHub repository with a commit hash. The seven benchmark datasets are cited third-party resources rather than paper-specific deposits, so only,Code · publicThe implementation is openly available at the GitHub repository https://github.com/
Fatema-Abdulqader/Agri-DSSA-Dual-Self-Supervised-Attention-Framework/tree/main, commit
98f3863Open asset ↗pdf-page:21 lines:1-61Plant phenotyping relevance match · UnverifiedCrossref · Europe PMC · checked 15 Sept 2026
Laboratory / benchtopMicroscopyMultimodalCell / cellular structureRootTissueTrackingVisualization / data management
Abstract Root biology is pivotal in addressing global challenges including sustainable agriculture and climate change. However, roots have been relatively understudied among plant organs, partly due to the difficulties in imaging root structures in their natural environment. Here we used microfabricated ecosystems (EcoFABs) to establish growing environments with optical access and employed nonlinear multimodal microscopy of third-harmonic generation (THG) and three-photon fluorescence (3PF) to achieve label-free, in situ imaging of live roots and microbes at high spatiotemporal resolution. THG enabled us to observe key plant root structures including the vasculature, Casparian strips, dividing meristematic cells, and root cap cells, as well as subcellular features including nuclear envelopes, nucleoli, starch granules, and putative stress granules. THG from the cell walls of bacteria and fungi also provides label-free contrast for visualizing these microbes in the root rhizosphere. With simultaneously recorded 3PF signal, we demonstrated our ability to investigate root-microbe interactions by achieving single-bacterium tracking and subcellular imaging of fungal spores and hyphae in the rhizosphere.
Why it matches plant phenotyping methodsTHG/3PFによる根の構造を高時空間分解能でラベルフリー取得するイメージング手法を開発・実証しており、植物表現型取得が中心である。
abstractemployed nonlinear multimodal microscopy of third-harmonic generation (THG) and three-photon fluorescence (3PF) to achieve label-free, in situ imaging of live roots and microbes at high spatiotemporal resolution
Early detection and diagnosis of plant diseases is critical for ensuring global food security and sustainable agricultural practices. This review comprehensively examines latest advancements in crop disease risk prediction, onset detection through imaging techniques, machine learning (ML), deep learning (DL), and edge computing technologies. Traditional disease detection methods, which rely on visual inspections, are time-consuming, and often inaccurate. While chemical analyses are accurate, they can be time consuming and leave less flexibility to promptly implement remedial actions. In contrast, modern techniques such as hyperspectral and multispectral imaging, thermal imaging, and fluorescence imaging, among others can provide non-invasive and highly accurate solutions for identifying plant diseases at early stages. The integration of ML and DL models, including convolutional neural networks (CNNs) and transfer learning, has significantly improved disease classification and severity assessment. Furthermore, edge computing and the Internet of Things (IoT) facilitate real-time disease monitoring by processing and communicating data directly in/from the field, reducing latency and reliance on in-house as well as centralized cloud computing. Despite these advancements, challenges remain in terms of multimodal dataset standardization, integration of individual technologies of sensing, data processing, communication, and decision-making to provide a complete end-to-end solution for practical implementations. In addition, robustness of such technologies in varying field conditions, and affordability has also not been reviewed. To this end, this review paper focuses on broad areas of sensing, computing, and communication systems to outline the transformative potential of end-to-end solutions for effective implementations towards crop disease management in modern agricultural systems. Foundation of this review also highlights critical potential for integrating AI-driven disease detection and predictive models capable of analyzing multimodal data of environmental factors such as temperature and humidity, as well as visible-range and thermal imagery information for early disease diagnosis and timely management. Future research should focus on developing autonomous end-to-end disease monitoring systems that incorporate these technologies, fostering comprehensive precision agriculture and sustainable crop production.
Why it matches plant phenotyping methods植物病害の画像・センサ計測、機械学習による病害分類・重症度推定を中心に扱うレビューであり、植物状態の取得・評価手法が主題。
abstractThis review comprehensively examines latest advancements in crop disease risk prediction, onset detection through imaging techniques, machine learning (ML), deep learning (DL), and edge computing technologies.
Brassica vegetablesLettuceRadishMicroscopyMultimodalRootVisualization / data management
Microfibers (MFs), primarily originating from sewage sludge and laundry effluents, are the most prevalent form of microplastics (MPs) in agricultural soils. While their ecological effects have been explored, the visualization, crop-level accumulation, and potential transport mechanisms of MFs within soil-plant systems remain poorly understood. This study combines 1,3,6,8-pyrene tetrasulfonic acid (PTSA) fluorescent staining with a sequential multimodal microscopy workflow to effectively track the distribution, adsorption, accumulation, and uptake of MFs under realistic soil cultivation conditions. Three edible vegetables-lettuce, Chinese cabbage, and cherry radish-were used to evaluate species-specific response patterns. The results revealed clear differences in MF interactions across species: lettuce exhibited strong MF adsorption on root surfaces and subsequent penetration via crack-entry and apoplastic pathways without entering cells. In contrast, Chinese cabbage and cherry radish showed limited MF adsorption and no uptake. These patterns were associated with root permeability and antioxidative capacities, indicating that plant functional traits play a critical role in determining the transport capacity of MPs. Beyond introducing a novel method for MF visualization in complex terrestrial matrices, this study provides new insights into the risks posed by MFs to soil-plant systems. The findings also highlight potential threats to food safety and underscore the need to establish plant-specific thresholds and pollution mitigation strategies to support sustainable agriculture and protect public health.
Why it matches plant phenotyping methods植物体内のマイクロファイバー分布・吸着・蓄積・取り込みを可視化する新規蛍光染色・マルチモーダル顕微鏡ワークフローが研究の中心であり、植物状態の測定法として該当する。
abstractThis study combines 1,3,6,8-pyrene tetrasulfonic acid (PTSA) fluorescent staining with a sequential multimodal microscopy workflow to effectively track the distribution, adsorption, accumulation, and uptake of MFs under realistic soil cultivation conditions.
Sweet potato (Ipomoea batatas L.) exhibits strong resilience in nutrient-poor soils and contains high levels of dietary fiber and antioxidant compounds. It also is highly tolerant to water stress, which has also contributed to its global distribution, particularly in regions prone to climatic variability. However, frequent abnormal climatic events have recently caused declines in both the quality and yield of sweet potatoes. To address this, machine learning (ML) and deep learning (DL) models based on a Vision Transformer-Convolutional Neural Network (ViT-CNN) were developed to classify water stress levels in sweet potato. RGB-thermal imagery captured from low-altitude platforms and various growth indicators were used to develop the classifier. The K-Nearest Neighbors (KNN) model outperformed other ML models in classifying water stress levels at all growth stages. The DL model simplified the original five-level water stress classification into three levels. This enhanced its sensitivity to extreme stress conditions, improve model performance, and increased its applicability to practical agricultural management strategies. To enhance practical applicability under open-field conditions, several environmental variables were newly defined to calculate the crop water stress index (CWSI). Furthermore, an integrated system was developed using gradient-weighted class activation mapping (Grad-CAM), explainable artificial intelligence (XAI), and a graphical user interface (GUI) to support intuitive interpretation and actionable decision-making. The system will be expanded into an online and fixed-camera platform to enhance its applicability to smart farming in diverse field crops.
Why it matches plant phenotyping methodsRGB・熱画像と生育指標を用いてサツマイモの水ストレス状態を分類するモデルを開発し、CWSI、XAI、GUIを統合したシステムを構築しており、植物状態の取得・推定手法が研究の中心である。
abstractmachine learning (ML) and deep learning (DL) models based on a Vision Transformer-Convolutional Neural Network (ViT-CNN) were developed to classify water stress levels in sweet potato.
The integration of unmanned aerial vehicles (UAVs) and deep learning (DL) has significantly advanced crop disease detection by enabling scalable, high-resolution, and near real-time monitoring within precision agriculture. This systematic review analyzes peer-reviewed literature indexed in the Web of Science Core Collection as articles or proceeding papers through 2024. The main selection criterion was combining “unmanned aerial vehicle*” OR “UAV” OR “drone” with “deep learning”, “agriculture” and “leaf disease” OR “crop disease”. Results show a marked surge in publications after 2019, with China, the United States, and India leading research contributions. Multirotor UAVs equipped with RGB sensors are predominantly used due to their affordability and spatial resolution, while hyperspectral imaging is gaining traction for its enhanced spectral diagnostic capability. Convolutional neural networks (CNNs), along with emerging transformer-based and hybrid models, demonstrate high detection performance, often achieving F1-scores above 95%. However, critical challenges persist, including limited annotated datasets for rare diseases, high computational costs of hyperspectral data processing, and the absence of standardized evaluation frameworks. Addressing these issues will require the development of lightweight DL architectures optimized for edge computing, improved multimodal data fusion techniques, and the creation of publicly available, annotated benchmark datasets. Advancements in these areas are vital for translating current research into practical, scalable solutions that support sustainable and data-driven agricultural practices worldwide.
Why it matches plant phenotyping methodsUAV画像と深層学習による作物病害の検出手法を体系的にレビューしており、植物の病害状態を観測・推定する方法論が中心である。
titleEvolution of Deep Learning Approaches in UAV-Based Crop Leaf Disease Detection: A Web of Science Review
The use of multiple camera technologies in a combined multimodal monitoring system for plant phenotyping offers promising benefits. Compared to configurations that rely on a single camera technology, cross-modal patterns can be recorded that allow a more comprehensive assessment of plant phenotypes. However, the effective utilization of cross-modal patterns depends on image registration to achieve pixel-precise alignment - a challenge often complicated by parallax and occlusion effects inherent in plant canopy imaging. In this study, we propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process. By leveraging depth data, our method mitigates parallax effects, facilitating more accurate pixel alignment across camera modalities. Additionally, we introduce an automated mechanism to identify and differentiate various types of occlusions, thereby minimizing registration errors. To evaluate the efficacy of our approach, we conduct experiments on a diverse dataset comprising six distinct plant species with varying leaf geometries. Our results demonstrate the robustness of the proposed registration algorithm, showcasing its ability to achieve accurate alignment across different plant types and camera compositions. Compared to previous methods our approach is not reliant on detecting plant-specific image features, making it suitable for a wide range of applications in plant sciences. Moreover, the registration approach can scale to arbitrary numbers of cameras with varying resolutions and wavelengths. Overall, our study contributes to advancing the field of plant phenotyping by offering a robust and reliable solution for multimodal image registration.
Why it matches plant phenotyping methods植物フェノタイピング向けのマルチモーダル3D画像位置合わせ手法を開発し、複数植物種のデータセットで性能評価しているため、方法が研究の中心である。
abstractIn this study, we propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process.
The use of multiple camera technologies in a combined multimodal monitoring system for plant phenotyping offers promising benefits. Compared to configurations that rely on a single camera technology, cross-modal patterns can be recorded that allow a more comprehensive assessment of plant phenotypes. However, the effective utilization of cross-modal patterns depends on image registration to achieve pixel-precise alignment - a challenge often complicated by parallax and occlusion effects inherent in plant canopy imaging. In this study, we propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process. By leveraging depth data, our method mitigates parallax effects, facilitating more accurate pixel alignment across camera modalities. Additionally, we introduce an automated mechanism to identify and differentiate various types of occlusions, thereby minimizing registration errors. To evaluate the efficacy of our approach, we conduct experiments on a diverse dataset comprising six distinct plant species with varying leaf geometries. Our results demonstrate the robustness of the proposed registration algorithm, showcasing its ability to achieve accurate alignment across different plant types and camera compositions. Compared to previous methods our approach is not reliant on detecting plant-specific image features, making it suitable for a wide range of applications in plant sciences. Moreover, the registration approach can scale to arbitrary numbers of cameras with varying resolutions and wavelengths. Overall, our study contributes to advancing the field of plant phenotyping by offering a robust and reliable solution for multimodal image registration.
Why it matches plant phenotyping methods植物フェノタイピング向けのマルチモーダル3D画像位置合わせ手法を開発・評価しており、表現型取得ワークフローの技術的中心である。
abstractwe propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process.
Typhoons are a primary cause of maize lodging, significantly reducing crop yield and resilience. Rapid and accurate assessment of lodging severity is essential for processing agricultural decision and implementing effective cropland management strategies. However, the crop lodging monitoring methods based on remote sensing require a large number of ground survey samples, which are time-consuming and labor-intensive, limiting their applicability for large-scale assessments during typhoon events. To address this challenge, this study proposed a Two-Step Augmentation Strategy (TSAS) that integrates Sentinel-2 and Sentinel-1 satellite data with limited field samples to automatically generate representative samples for different lodging severity. The TSAS framework includes two steps: first step, we use eXtreme Gradient Boosting (XGBoost) to classify random points in overlapping regions (RegionOₗ, referring to areas covered by both Sentinel-1 and Sentinel-2 images) based on lodging-sensitive optical and radar features identified through correlation analysis and J-M distance, generating Sample Set L. Second step, we retrain the XGBoost model with radar features from Sample Set L to classify random points in non-overlapping regions (RegionNₒₗ, referring to areas covered only by Sentinel-1 images and not by Sentinel-2 images). Finally, a purification strategy is applied to remove erroneous samples and improve accuracy. This approach was tested in Jilin Province, significantly affected by a severe typhoon in 2020. Results show that TSAS effectively generates reliable lodging samples, achieving high spectral correlation similarity (mean of SCS > 0.72) and low Euclidean distance (mean of ED < 0.75) compared to field survey samples. These samples were applied to three classifiers, with Support Vector Machine (SVM) achieving the highest accuracy (OA = 82.36 %, F1 = 0.81) compared to Random Forest (RF) and Minimum Distance (MD) classification methods. The results of this study show that TSAS is able to efficiently and accurately generate high-accuracy maize lodging samples on a regional scale with limited field surveys, significantly improves the efficiency of large-scale agricultural disaster assessment and provides valuable support for disaster management and agricultural planning.
Why it matches plant phenotyping methodsSentinel-1/2データとXGBoostを用いて、トウモロコシの倒伏重症度サンプルを限られた現地調査から自動生成するTSAS手法が研究の中心であり、植物状態の推定と精度評価を実施している。
abstractthis study proposed a Two-Step Augmentation Strategy (TSAS) that integrates Sentinel-2 and Sentinel-1 satellite data with limited field samples to automatically generate representative samples for different lodging severity.
This paper presents a robotic system designed to replace manual labour in three prevalent activities related to blueberry management: soil sampling and analysis, weed spraying, and plant health status monitoring. A complex system for automating field activities in blueberry orchards that involves the use of ground and aerial robots, along with integrated route optimisation software, autonomous driving, image recognition based on AI, and others, was developed. The ground robot is made in a modular way, having three different add-on modules for the three distinct use cases it covers. A modular system guarantees the year-round utilisation of the robot across various growth stages of plants. This stands out as its major advantage, considering that many robotic systems are typically tailored to address a singular task. Robot modules utilise custom-made hardware for in-field sampling and real-time soil analysis, an industrial robotic arm with a custom-made system for spraying, and a plant health monitoring module consisting of two Plant-O-Meter™ optical devices capable of capturing sixteen optical vegetation indices in real-time. The solutions are tested and deployed in the real-world environment of the blueberry orchard. We have achieved 50–60 % savings on herbicides compared to blanket spraying, and the costs related to sub-optimal herbicide application are reduced up to 25 %. The costs of soil analysis have been significantly reduced to about 70 %.
Why it matches plant phenotyping methodsブルーベリー園向けロボット基盤の主要モジュールとして、植物健康状態を光学センサーでリアルタイム測定する方法を実環境で展開・評価しており、植物状態の取得が技術的構成の中心的要素の一つである。
abstracta plant health monitoring module consisting of two Plant-O-Meter™ optical devices capable of capturing sixteen optical vegetation indices in real-time
Abstract Accurate identification of maize diseases is crucial for safeguarding global food security. Traditional image-based methods often struggle with lighting variations, occlusions, and noise, which limits their robustness and generalisation ability. Multimodal approaches that integrate visual and textual information have shown promise. However, these methods frequently require manually curated textual descriptions for each image, increasing data collection costs and limiting scalability and practical implementation. To address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA). This approach enforces category-level alignment between image and text modalities, enabling more accurate and interpretable cross-modal mapping. First, we construct cross-modal representations by aligning image and text modalities at the category level within a shared embedding space. Second, inspired by contrastive learning, we introduce a Cross-Modal Category Alignment (CMCA) loss based on category-level textual descriptions, reducing annotation complexity. Finally, we present an Efficient Channel-Spatial Hybrid Attention (CSHA) module that preserves inter-class boundaries with minimal computational overhead to enhance feature discriminability under complex conditions. Experimental results show that the mIT-CMCA model achieves accuracy, precision, recall, and F1-scores of 99.48%, 99.28%, 99.54%, and 99.41%, respectively, on the maize subset of the PlantVillage dataset (MPVD), representing improvements of 0.24%, 0.13%, 0.17%, and 0.15% over the strongest baseline. On the self-built Maize Leaf-Field dataset (MLFD), the model attains corresponding scores of 93.67%, 93.76%, 93.67%, and 93.66%, with respective improvements of 5.34%, 5.13%, 5.34%, and 5.38%.
Why it matches plant phenotyping methodsトウモロコシ葉画像から病害を識別する画像・テキスト解析手法を開発し、複数データセットで性能評価している。病害状態の画像ベース推定が研究の中心である。
abstractTo address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA).
Accurate plant segmentation in thermal imagery remains a significant challenge for high throughput field phenotyping, particularly in outdoor environments where low contrast between plants and weeds and frequent occlusions hinder performance. To address this, we present a framework that leverages synthetic RGB imagery, a limited set of real annotations, and GAN-based cross-modality alignment to enhance semantic segmentation in thermal images. We trained models on 1,128 synthetic images containing complex mixtures of crop and weed plants in order to generate image segmentation masks for crop and weed plants. We additionally evaluated the benefit of integrating as few as five real, manually segmented field images within the training process using various sampling strategies. When combining all the synthetic images with a few labeled real images, we observed a maximum relative improvement of 22% for the weed class and 17% for the plant class compared to the full real-data baseline. Cross-modal alignment was enabled by translating RGB to thermal using CycleGAN-turbo, allowing robust template matching without calibration. Results demonstrated that combining synthetic data with limited manual annotations and cross-domain translation via generative models can significantly boost segmentation performance in complex field environments for multi-model imagery.
Why it matches plant phenotyping methods熱画像による作物・雑草の分割を対象とし、合成データ、少数の実画像、GANによるモダリティ変換を組み合わせた高スループット表現型取得手法を開発・評価している。
abstractAccurate plant segmentation in thermal imagery remains a significant challenge for high throughput field phenotyping
Reproduction assets foundThe paper states its synthetic and real phenotyping image datasets (cowpea/weed RGB and thermal imagery with segmentation masks) are publicly available through the AgML framework, with an explicit authors' URL. Helios is a general simulation tool, not a paper-specific asset.Dataset · publicSynthetic and real datasets are available through AgML 1 1
1
https://github.com/Project-AgML/AgML [ 50 ] , a centralized framework for agricultural machine learning.Open asset ↗Project-AgML/AgMLlines:339-434Plant phenotyping relevance match · UnverifiedCrossref · checked 13 Sept 2026
Forest aboveground biomass (AGB) is a key component of terrestrial carbon storage, essential for understanding the carbon cycle and evaluating carbon sink potential. However, estimating long-term AGB in tropical forests and detecting its spatial and temporal trends remain challenging due to observational gaps and methodological constraints. Here, we integrate GEDI L4B gridded biomass data with features from MODIS, PALSAR/PALSAR-2, SRTM, and climate datasets, and apply the AutoGluon ensemble learning framework to develop AGB retrieval models. We generated annual AGB maps at 1 km resolution for Borneo’s forests from 2007 to 2023, achieving high predictive accuracy (R2 = 0.92, RMSE = 32.84 Mg/ha, rRMSE = 21.06%). Residuals were generally balanced and close to a symmetric distribution, indicating no strong bias within the moderate biomass range (50–350 Mg/ha). However, in very high-biomass stands, the model tended to underestimate AGB, reflecting saturation effects that persist despite clear improvements over existing products. Estimated mean AGB values ranged from 180.52 to 214.09 Mg/ha, with total AGB varying between 13.05 and 14.10 Pg. Trend analysis using Sen’s slope and the Mann–Kendall test revealed significant AGB trends in 31.31% of forested areas, with 68.76% showing increases. This study offers a robust and scalable framework for continuous tropical forest carbon monitoring, providing critical support for carbon accounting, forest management, and policy-making.
Why it matches plant phenotyping methods衛星・LiDAR等のマルチセンサーデータと機械学習により、森林の地上部バイオマスという植物群落形質を推定する手法を開発・検証しており、形質推定法が中心である。
abstractwe integrate GEDI L4B gridded biomass data with features from MODIS, PALSAR/PALSAR-2, SRTM, and climate datasets, and apply the AutoGluon ensemble learning framework to develop AGB retrieval models.
Street trees are vital to urban livability, providing ecological and social benefits. Establishing a detailed, accurate, and dynamically updated street tree inventory has become essential for optimizing these multifunctional assets within space-constrained urban environments. Given that traditional field surveys are time-consuming and labor-intensive, automated surveys utilizing Mobile Mapping Systems (MMS) offer a more efficient solution. However, existing MMS-acquired tree datasets are limited by small-scale scene, limited annotation, or single modality, restricting their utility for comprehensive analysis. To address these limitations, we introduce WHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset. Collected across two distinct cities, WHU-STree integrates synchronized point clouds and high-resolution images, encompassing 21,007 annotated tree instances across 50 species and 2 morphological parameters. Leveraging the unique characteristics, WHU-STree concurrently supports over 10 tasks related to street tree inventory. We benchmark representative baselines for two key tasks--tree species classification and individual tree segmentation. Extensive experiments and in-depth analysis demonstrate the significant potential of multi-modal data fusion and underscore cross-domain applicability as a critical prerequisite for practical algorithm deployment. In particular, we identify key challenges and outline potential future works for fully exploiting WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV/WHU-STree.
Why it matches plant phenotyping methods樹木の個体セグメンテーションと形態パラメータを含むマルチモーダルデータセットを構築し、ベンチマークする研究であり、植物個体の状態・形態抽出手法が中心である。
abstractWHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset.
Reproduction assets foundThe paper's core asset is the WHU-STree multi-modal street tree dataset (point clouds, panoramic images, 21,007 annotated tree instances, 50 species, height/DBH), which the authors state is publicly accessible via their GitHub organization WHU-USI3DV. The Zenodo DOIs in the reference list belong to cited prior datasetsDataset · publicticular, we
identify key challenges and outline potential future works for fully exploit-
ing WHU-STree, encompassing multi-modal fusion, multi-task collaboration,
cross-domain generalization, spatial pattern learning, and Multi-modal Large
Language Model for street tree asset management. The WHU-STree dataset
is accessible at: https://github.com/WHU-USI3DV /WHU-STree.
Keywords: Deep learning, Tree inventory, Individual tree segmentation,
Tree species classification, Multi-modal, Mobile mapping system
1. Introduction
Street trees, vital to urban ecosystems, provide ecological benefits (e.g.,
shade (Kumar et al., 2024), air purification (Grundstrém and Pleijel, 2014),
noise reductiOpen asset ↗WHU-STreepdf-raw-page:2 lines:1-35Plant phenotyping relevance match · UnverifiedEurope PMC · Crossref · checked 14 Sept 2026
Sustainable agriculture urgently requires innovative, pesticide-free strategies to mitigate herbivory and safeguard food security. Ultraviolet-B (UV-B) irradiation, with tunable intensity and cost-effectiveness, has emerged as a promising non-chemical method to enhance plant resistance, yet its underlying mechanisms remain elusive. Here, using tea plant ( Camellia sinensis ) and its major pest Ectropis obliqua as a model, we developed a multimodal framework that integrates AI-enhanced electronic nose technology for real-time volatile profiling with in situ hyperspectral stimulated Raman scattering (SRS) microscopy to characterize defense responses under precisely controlled UV-B treatments. This approach identified herbivore-induced volatiles—hexanal, (Z)-3-hexenol, octanal, and (Z)-3-hexenyl acetate—optimally induced at 1.2 kJ·m -2 UV-B and linked to insect deterrence. SRS imaging further revealed elevated jasmonic acid derivatives and L-phenylalanine, coupled with reduced protein levels and altered stomatal dynamics, all correlating with enhanced resistance. Transcriptomic and molecular analyses confirmed transcriptional regulation of these pathways. By bridging volatile detection, metabolic imaging, and molecular validation, this study pioneers a multimodal strategy that provides mechanistic insights into UV-B–mediated plant defense and highlights the potential of multimodal methodologies as powerful tools for developing sustainable, pesticide-free pest management solutions in precision agriculture.
Why it matches plant phenotyping methodsAI強化電子鼻とSRS顕微鏡を統合した植物防御応答の取得・解析フレームワークが研究の中心であり、揮発性物質、代謝、気孔動態などの植物状態を測定しているため。
abstractwe developed a multimodal framework that integrates AI-enhanced electronic nose technology for real-time volatile profiling with in situ hyperspectral stimulated Raman scattering (SRS) microscopy to characterize defense responses
Crop diseases remain a threat to the world's food security with yield loss ranging from 10–40% annually. The last few years have witnessed spectacular evolution in artificial intelligence (AI), deep learning, Internet of Things (IoT), and unmanned aerial vehicles (UAVs), which transformed crop disease monitoring and early detection of diseases. Earlier image processing methods are now overpowered by convolutional neural networks (CNNs), object detectors such as the YOLO family, and CNN-transformer hybrids, which are significantly more accurate and robust. At the same time, IoT sensors and UAV-based multispectral imaging provide complementary environmental and spectral information for enabling active monitoring irrespective of visible indicators. But there are some limitations they have, which are poor model generalization when trained from human-annotated datasets, expensive computation in field deployment, lack of rich plentiful annotated data for low-frequency diseases, and challenging adoption in smallholder farming settings.Here, the three areas of crop health monitoring system development are critically evaluated as follows: (i) image-based systems, (ii) deep learning models, and (iii) multimodal integration of UAV and IoT. Critical comparative performance, strength, and weakness of current methods are analyzed, highlighting dataset heterogeneity, detection accuracy, scalability, and practicability of deployment. Additionally, the review reveals the key deficits in the research—i.e., necessity for robust multimodal fusion paradigms, conventional benchmarking, and affordable field solutions—and suggests likely future directions such as federated learning, predictive outbreak modeling, and robotics for targeted intervention. Synthesizing current success and pointing toward likely future research directions, this review seeks to inform researchers and practitioners toward sustainable, tech-enabled crop disease management.
Why it matches plant phenotyping methods作物の病害状態を画像・深層学習・UAV・IoTセンサーで検出・監視する方法を中心に比較評価したレビューであり、植物フェノタイピング手法のレビューに該当する。
titleDeep Learning Approaches for Crop Health Monitoring and Early Disease Detection: A Review
Tomato plants, a globally significant horticultural crop, are frequently threatened by a range of diseases that compromise yield and quality. Traditional disease detection methods based on manual inspection by experts are often labour-intensive, time-consuming, and susceptible to human error. This study presents a machine learning-based approach that leverages Convolutional Neural Networks (CNNs) to automate the identification and classification of common tomato diseases using leaf images. A comprehensive dataset, including both healthy and diseased leaf images, was collected, pre-processed, and augmented to enhance model performance under various environmental conditions. A custom-designed CNN model was then trained and evaluated using standard metrics such as accuracy, precision, recall, and F1-score. The model demonstrated high classification accuracy and robustness across multiple disease categories including early blight, late blight, and bacterial spot. Furthermore, the system was deployed as a user-friendly web and mobile application interface, allowing real-time diagnosis in the field. This enables farmers especially in resource-constrained settings to identify and respond to infections early, thereby reducing yield losses and limiting excessive pesticide use. The project underscores the potential of AI-driven solutions in modernizing agricultural practices and promoting sustainable crop management. Recommendations are made for future work to improve model adaptability, extend its disease coverage, and integrate environmental sensor data for multimodal analysis.
Why it matches plant phenotyping methodsトマト葉画像から病害状態を分類するCNN手法を開発・評価しており、植物の病徴・病害状態の取得が研究の中心であるため。
abstractThis study presents a machine learning-based approach that leverages Convolutional Neural Networks (CNNs) to automate the identification and classification of common tomato diseases using leaf images.
Poplar trees are widely cultivated for their ecological and economic benefits. Studying the phenotypes of poplar seedlings can enable the selection of optimal cultivation methods to enhance yield and quality. UAV-based low-altitude remote sensing with optical sensors captures images and spectral data for such studies. However, deep learning in UAV plant phenotyping faces the challenge of requiring substantial time and effort to label image samples for model training. This paper aims to assess the efficiency of using Grounding DINO-SAM2 for zero-shot instance segmentation of individual poplar seedlings across multiple genotypes. An automatic program calculates image features from RGB and multispectral mask areas, including canopy projection, color, texture, and spectral reflectance, which are then used to establish a biomass estimation model based on two years of data. The study obtained the following results: (1) The Grounding DINO-SAM2 model was used to implement zero-labelled sample instance segmentation of 400 image data. After modifying the sample with incorrect target recognition quantity in less than 15 min, the total model took only 0.5 h, with a precision of 0.943, which greatly saved time and computing cost compared with mainstream fully-supervised segmentation models. (2) A poplar seedling biomass estimation model based on multimodal image features was established. After comparing and optimizing single-sensor and multi-sensor combined with different modelling algorithms, it was found that the CNN test set accuracy (R²) reached 0.823. This research provides a lightweight, cost-effective approach for plant image segmentation and feature extraction, promoting advances in intelligent management and monitoring for agriculture and forestry.
Why it matches plant phenotyping methodsUAV画像による個体セグメンテーションと特徴抽出を開発・評価し、ポプラ苗のバイオマスを推定する方法が研究の中心である。
abstractThis paper aims to assess the efficiency of using Grounding DINO-SAM2 for zero-shot instance segmentation of individual poplar seedlings across multiple genotypes.
Anthesis prediction is crucial for breeding wheat. While current tools provide estimates of average anthesis at the field scale, they fail to address the needs of breeders who require accurate predictions for individual plants. Hybrid breeders have to finalize their plans for pollination at least 10 days before such flowering is due and biotechnology field trials in the United States and Australia must report to regulators 7–14 days before the first plant flowers. Currently, predicting anthesis of individual wheat plants is a labour-intensive, inefficient, and costly process. Individual wheat of the same cultivar within the same field may exhibit substantial variations in anthesis timing, due to significant variations in their immediate surroundings. In this study, we developed an efficient and cost-effective machine vision approach to predict anthesis of individual wheat plants. By integrating RGB imagery with in-situ meteorological data, our multimodal framework simplifies the anthesis prediction problem into binary or three-class classification tasks, aligning with breeders' requirements in individual wheat flowering prediction on the crucial days before anthesis. Furthermore, we incorporated a few-shot learning method to improve the model's adaptability across different growth environments and to address the challenge of limited training data. The model achieved an F1 score above 0.8 in all planting settings.
Why it matches plant phenotyping methods個体コムギの開花(anthesis)時期という植物形質を、RGB画像・気象データ・few-shot学習で予測する機械視覚手法の開発が中心であり、フェノタイピング手法として採用。
abstractIn this study, we developed an efficient and cost-effective machine vision approach to predict anthesis of individual wheat plants.
Rapid and accurate monitoring of crop water status is essential for ensuring sustainable agricultural development and food security. Crops exhibit a complex set of growth and physiological responses under water deficit. Existing studies primarily focused on the monitoring of phenotypic parameters, while the physiological indicators highly relevant to crop water status were ignored. In this context, we aimed to develop a novel model to comprehensively and accurately monitor maize water status using multi-source UAV data and multiple growth and physiological indicators. We first composed the original dataset, including feature variables based on multi-source UAV data (spectral indices, texture indices, thermal indices, and structural indices) and prediction variables based on field measurements (equivalent water thickness, stomatal conductance, transpiration rate, and actual photochemical efficiency) in 2023 and 2024. Next, the tabular denoising diffusion probabilistic model (TabDDPM) was employed for synthesizing new samples to adequately train the models. Then, a deep learning network named TAM-Net, with the hybrid attention mechanism and multi-task learning, was trained on the synthetic dataset. Finally, the fuzzy comprehensive water index (FCWI) considering uncertainty and variability was obtained in 2023–2024. The results indicated that multi-source data significantly improved the model performance, with the R² of 0.52–0.63, and NRMSE of 27.89 %–29.86 %. TabDDPM was able to synthesize new datasets with high similarity and effectiveness. TAM-Net achieved the highest monitoring accuracy for the four indicators of crop water status (R² of 0.76–0.90, NRMSE of 12.92 %–22.95 %). FCWI effectively assessed the water status across different treatments. Overall, TAM-Net was demonstrated with powerful performance for monitoring maize water status, which has potential in supporting precision irrigation practices.
Why it matches plant phenotyping methodsUAVマルチソース画像からトウモロコシの水分状態・生理形質を推定するTAM-Netを開発し、精度評価まで行っており、表現型取得・推定手法が研究の中心である。
abstractwe aimed to develop a novel model to comprehensively and accurately monitor maize water status using multi-source UAV data and multiple growth and physiological indicators.
High-throughput plant phenotyping (HTP) has emerged as a crucial element of modern plant science and crop breeding, enabling the swift, large-scale, and non-invasive evaluation of plant characteristics. The integration of machine learning (ML), particularly deep learning approaches, has transformed HTP by automating feature extraction, improving predictive accuracy, and enabling thorough analysis under diverse environmental conditions. This review compiles the latest advancements in ML-based phenotyping, covering traditional algorithms, convolutional neural networks, transformer models, and innovative 3D reconstruction techniques. It investigates advanced phenotyping technologies, encompassing controlledenvironment systems, field robotics, and UAV-driven imaging, along with novel instruments such as ChronoRoot 2.0 and PhenoAssistant. The conversation includes uses in yield forecasting, stress identification, trait measurement, and weed differentiation. Additionally, it examines significant issues like dataset constraints, interpretability of models, and scalability, while suggesting future paths that involve multimodal integration, open data standards, explainable AI, and affordable phenotyping methods. This synthesis is designed to help researchers leverage ML for phenotyping processes, thereby promoting precision agriculture and accelerating breeding initiatives
Why it matches plant phenotyping methods植物フェノタイピングにおける機械学習手法、画像・3D再構成、ロボティクス、UAV、データセット課題などを中心に扱う方法論レビューである。
titleAdvances in Machine Learning for High-Throughput Plant Phenotyping: Techniques, Applications, and Opportunities
Rapid and accurate quantification of mineral elements in plants facilitates the optimization of cultivation strategies and provides theoretical support for heavy metal pollution control. Compared to traditional chemical detection methods, laser-induced breakdown spectroscopy (LIBS) offers rapid, simultaneous multi-element analysis. However, the quantitative accuracy of LIBS is often hindered by challenges such as sample heterogeneity and the inherent matrix effects arising from the physical and chemical properties of samples. These limitations highlight the need for innovative approaches to improve the reliability and precision of LIBS-based elemental quantification. In this study, we proposed a low-cost image-spectroscopy dual-modal rapid detection system combined with a dual-modal hierarchical fusion network (DMH-FNet). Compared with a standalone LIBS system, the quantification performance improved for the seven elements, namely P, Ca, Mg, Zn, Mn, K, and Si. During the validation phase, feature map visualization was used to interpret the feature extraction process of DMH-FNet. The results indicate that the model shifted its focus from low-level features of ablation crater details to high-level global features of the sample. Subsequently, SHapley additive exPlanations (SHAP) was used to explain the decision-making process of the optimal quantitative model and visualize key image features. The results demonstrate that DMH-FNet efficiently extracts features highly correlated with ablation crater information using its neural network capabilities and enhances the quantification of mineral elements through complementary fusion with LIBS spectral features. This study is the first to leverage the superior feature extraction capability of neural networks to capture valuable information from ablation images, thereby improving the quantification performance for multiple mineral elements. In conclusion, the proposed detection system, with its low-cost equipment, real-time data acquisition, and DMH-FNet, enables simultaneous, rapid, and accurate prediction of multiple mineral element contents.
Why it matches plant phenotyping methods植物葉の鉱元素含量という生理形質を対象に、LIBSと画像を融合した低コスト検出システムおよびニューラルネットワークを開発・検証しており、形質取得法が研究の中心である。
abstractwe proposed a low-cost image-spectroscopy dual-modal rapid detection system combined with a dual-modal hierarchical fusion network (DMH-FNet).
Digital twin applications offered transformative potential by enabling real-time monitoring and robotic simulation through accurate virtual replicas of physical assets. The key to these systems is 3D reconstruction with high geometrical fidelity. However, existing methods struggled under field conditions, especially with sparse and occluded views. This study developed a two-stage framework (DATR) for the reconstruction of apple trees from sparse views. The first stage leverages onboard sensors and foundation models to semi-automatically generate tree masks from complex field images. Tree masks are used to filter out background information in multi-modal data for the single-image-to-3D reconstruction at the second stage. This stage consists of a diffusion model and a large reconstruction model for respective multi view and implicit neural field generation. The training of the diffusion model and LRM was achieved by using realistic synthetic apple trees generated by a Real2Sim data generator. The framework was evaluated on both field and synthetic datasets. The field dataset includes six apple trees with field-measured ground truth, while the synthetic dataset featured structurally diverse trees. Evaluation results showed that our DATR framework outperformed existing 3D reconstruction methods across both datasets and achieved domain-trait estimation comparable to industrial-grade stationary laser scanners while improving the throughput by $\sim$360 times, demonstrating strong potential for scalable agricultural digital twin systems.
Why it matches plant phenotyping methodsリンゴ樹の疎視点画像から3D形状を再構成し、樹体形質を推定する手法の開発と、圃場・合成データでの評価が研究の中心であるため。
abstractThis study developed a two-stage framework (DATR) for the reconstruction of apple trees from sparse views.
The rapid development of intelligent technologies has transformed various industries, and agriculture benefits greatly from precision farming innovations. One of the remarkable achievements in agriculture is enhancing pest and disease identification for better crop health control and higher yields. This paper presents novel models of a multimodal data fusion technique to meet the growing need for accurate and timely wheat pest and disease identification. It combines image processing, sensor - derived environmental data, and machine learning for reliable wheat pest and disease diagnosis. First, deep - learning algorithms in image analysis detect early - stage pests and diseases on wheat leaves. Second, environmental data such as temperature and humidity improve diagnosis. Third, the data fusion process integrates image data for further analysis. Finally, several criteria compare the proposed model with previous methods. Experimental results show the proposed techniques achieve a detection accuracy of 96.5%, precision of 94.8%, recall of 97.2%, F1 score of 95.9%, MCC of 0.91, and AUC - ROC of 98.4%. The training time is 15.3 hours, and the inference time is 180 ms. Compared with CNN - based and SVM - based techniques, the proposed model's improvement is analyzed. It can be adapted for real - time use and applied to more crops and diseases.
Why it matches plant phenotyping methods小麦葉の病害・害虫状態を画像と環境センサーデータから推定するマルチモーダル手法の開発・比較が中心であり、植物病害フェノタイピングに該当する。
abstractThis paper presents novel models of a multimodal data fusion technique to meet the growing need for accurate and timely wheat pest and disease identification.
With the advancement of remote sensing imagery and multimodal sensing technologies, monitoring plant trait dynamics has emerged as a critical area of research in modern agriculture. Traditional approaches, which rely on handcrafted features and shallow models, struggle to effectively address the complexity inherent in high-dimensional and multisource data. In contrast, deep learning, with its end-to-end feature extraction and nonlinear modeling capabilities, has substantially improved monitoring accuracy and automation. This review summarizes recent developments in the application of deep learning methods—including CNNs, RNNs, LSTMs, Transformers, GANs, and VAEs—to tasks such as growth monitoring, yield prediction, pest and disease identification, and phenotypic analysis. It further examines prominent research themes, including multimodal data fusion, transfer learning, and model interpretability. Additionally, it discusses key challenges related to data scarcity, model generalization, and real-world deployment. Finally, the review outlines prospective directions for future research, aiming to inform the integration of deep learning with phenomics and intelligent IoT systems and to advance plant monitoring toward greater intelligence and high-throughput capabilities.
Why it matches plant phenotyping methods植物形質モニタリングとフェノタイピングにおける深層学習手法を体系的に扱うレビューであり、方法論が中心です。
abstractThis review summarizes recent developments in the application of deep learning methods—including CNNs, RNNs, LSTMs, Transformers, GANs, and VAEs—to tasks such as growth monitoring, yield prediction, pest and disease identification, and phenotypic analysis.
The biodiversity function of the desert steppe ecosystem faces many challenges under the pressure of climate change and human activities. Accurate and efficient assessment of plant diversity is critical for guiding desert steppe restoration efforts. However, desert steppe vegetation has sparse leaves and sparse distribution. It is difficult to accurately distinguish micro-vegetation types based on a single spectrum, vegetation index or texture feature, and the resolution of satellite remote sensing cannot meet the needs of high-precision diversity assessment. To this end, this study proposed a novel method for assessing plant diversity index in degraded desert grassland based on multimodal UAV hyperspectral data and Encoder-CNN. Through experiments on different modal feature combinations, spatial spectra, vegetation indices and texture features were targeted and fused. Channel Attention Fusion (CAF) was introduced into Encoder to achieve cross-layer "soft" residual fusion, the Encoder and CNN models were fused to construct a global-local co-expression structure, and finally the quantitative calculation of the plant diversity index at the pixel level was realized. The results show that the vegetation types determined by the fusion of multimodal data and deep learning are consistent with the existing species, dominant species and sub-dominant species of the actual community, and the calculated diversity index results are also consistent with the actual situation. The use of multimodal data combining spatial spectral features with index features, combined with the Encode-CNN model, can provide the most accurate information on community composition. The overall accuracy of sparse vegetation classification can reach 90.01%, and the average accuracy can reach 85.23%, which is better than single mode or traditional 3DCNN, VIT models. This study demonstrates the application potential of UAV hyperspectral multimodal technology and deep learning in the assessment of desert steppe plant diversity, providing important technical support for ecological protection and conservation.
Why it matches plant phenotyping methodsUAVハイパースペクトルとEncoder-CNNを用いて、植物多様性指数を画素レベルで定量推定する手法を開発・評価しており、植物状態の取得・抽出が研究の中心である。
abstractthis study proposed a novel method for assessing plant diversity index in degraded desert grassland based on multimodal UAV hyperspectral data and Encoder-CNN.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。Code · publicThe codes used in this study are available at https://github.com/15204718180/encoder-cnn.Open asset ↗15204718180/encoder-cnnpdf-page:17 lines:56-74Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
Published20 Aug 2025International Journal of Latest Technology in Engineering Management & Applied ScienceCited by 1 · OpenAlex ↗
Abstract. Plant diseases are a critical barrier to agricultural sustainability, contributing to annual crop losses of 30–40% in some regions. Such diseases arise from diverse pathogens, including fungi, bacteria, and viruses, and can severely degrade crop yield and quality if undetected. Early diagnosis is crucial for timely intervention, reduced pesticide use, and long-term soil health. Traditional approaches, which rely on manual visual inspection of leaves, remain slow, subjective, and impractical for monitoring large or remote farms. In recent years, the convergence of artificial intelligence (AI), computer vision, and image-based analysis has enabled automated disease detection systems that are both scalable and real-time. Classical techniques first employed color, texture, and shape descriptors combined with machine learning models such as Support Vector Machines (SVM), k-Nearest Neighbor (KNN), and Random Forests. These methods achieved moderate success but struggled with variability in lighting, background, and leaf orientation. The introduction of deep learning, particularly Convolutional Neural Networks (CNNs), transformed plant disease detection by allowing end-to-end feature learning directly from raw images. Modern lightweight architectures like MobileNet and EfficientNet further enhance deployability on mobile and edge devices. In parallel, segmentation-assisted and hybrid models improve robustness under complex field conditions. This review consolidates these developments, evaluating methods in terms of accuracy, generalization, computational efficiency, and field-readiness. It also identifies persistent challenges—such as limited annotated datasets and the opaque decision-making of deep models—and highlights future directions in explainable AI, multi-modal sensing, and IoT-integrated precision agriculture.
Why it matches plant phenotyping methods植物病害を画像から検出するAI・画像処理手法を比較評価するレビューであり、植物状態の表現型取得法が中心です。
titleFeature Engineering to Early Detection of Plant Disease Using Image Processing and Artificial Intelligence: A Comparative Analysis
Global agricultural sustainability and food security are increasingly challenged by pervasive crop diseases, necessitating advanced, real-time detection systems. Existing methodologies often struggle with the limitations of visual symptom reliance, generalizability across diverse crop types, and integration into continuous monitoring paradigms. This paper proposes a comprehensive framework for sustainable crop health monitoring that integrates cutting-edge Convolutional Neural Network (CNN) architectures with multimodal sensing techniques. We highlight the strategic integration of signals from diverse sources, such as RGB, thermal, and hyperspectral imagery, to record subtle physiological biomarkers that signal early disease onset, transcending traditional visual diagnostics. Our experimental validations establish the superiority of the framework's performance at a remarkable 96.8% accuracy and an important 91.3% early disease detection rate. This far surpasses current unimodal methods, which generally have an early detection rate of just 65.2% for a standard RGB CNN. By promoting an astute, evidence-based method to crop disease management, this work makes a valuable contribution to increasing agricultural resilience, improving resource efficiency, and promoting environmentally conscious agriculture globally. It demonstrates its effectiveness in detecting diseases at their earliest stages.
Why it matches plant phenotyping methodsCNNとRGB・熱・ハイパースペクトル画像を用いて植物病害状態を検出する方法の提案・検証が中心であり、植物表現型として病害状態を直接推定している。
abstractThis paper proposes a comprehensive framework for sustainable crop health monitoring that integrates cutting-edge Convolutional Neural Network (CNN) architectures with multimodal sensing techniques.
High-quality disease segmentation plays a crucial role in the precise identification of rice diseases. Although the existing deep learning methods can identify the disease on rice leaves to a certain extent, these methods often face challenges in dealing with multi-scale disease spots and irregularly growing disease spots. In order to solve the challenges of rice leaf disease segmentation, we propose KBNet, a novel multi-modal framework integrating language and visual features for rice disease segmentation, leveraging the complementary strengths of CNN and Transformer architectures. Firstly, we propose the Kalman Filter Enhanced Kolmogorov-Arnold Networks (KF-KAN) module, which combines the modeling ability of KANs for nonlinear features and the dynamic update mechanism of the Kalman filter to achieve accurate extraction and fusion of multi-scale lesion information. Secondly, we introduce the Boundary-Constrained Physical-Information Neural Network (BC-PINN) module, which embeds the physical priors, such as the growth law of the lesion, into the loss function to strengthen the modeling of irregular lesions. At the same time, through the boundary punishment mechanism, the accuracy of edge segmentation is further improved and the overall segmentation effect is optimized. The experimental results show that the KBNet framework demonstrates solid performance in handling complex and diverse rice disease segmentation tasks and provides key technical support for disease identification, prevention, and control in intelligent agriculture. This method has good popularization value and broad application potential in agricultural intelligent monitoring and management.
Why it matches plant phenotyping methodsイネ葉の病斑という植物の病態を画像からセグメンテーションする新規手法を開発しており、病害識別のための表現抽出・境界推定が研究の中心である。
abstractwe propose KBNet, a novel multi-modal framework integrating language and visual features for rice disease segmentation
Reproduction assets foundThe paper's rice disease segmentation images (1550 leaf images of Blast, Bacterial Blight, and Tungro) were sourced from the public PlantVillage dataset on Kaggle, which is publicly available at the allowed URL. However, the paper-specific derived assets — the authors' pixel-level Labelme segmentation masks, text描述, 4:Dataset · publici: 10.1007/s10462-024-10717-2.
26. Li A., Jiao L., Zhu H., Li L., Liu F. Multitask Semantic Boundary Awareness Network for Remote Sensing Image Segmentation. IEEE Trans. Geosci. Remote Sens. 2022;60:1–14. doi: 10.1109/TGRS.2021.3050885.
27. Kaggle, PlantVillage Dataset. 2019. [(accessed on 19 September 2022)]. Available online: https://www.kaggle.com/datasets/abdallahalidev/plantvillage-dataset .
28. Li Z., Li Y., Li Q., Wang P., Guo D., Lu L., Jin D., Zhang Y., Hong Q. LViT: Language Meets Vision Transformer in Medical Image Segmentation. IEEE Trans. Med. Imaging. 2024;43:96–107. doi: 10.1109/TMI.2023.3291719.
29. Liu Z., Wang Y., Vaidya S., Ruehle F., Halverson J., Soljacic M., Hou T.Y., TOpen asset ↗Kaggle · plantvillage-datasetlines:409-431Plant phenotyping relevance match · UnverifiedEurope PMC · checked 15 Sept 2026
Published1 Aug 2025Computers and Electronics in Agriculture.
Dwarf tomatoes, with high edible and ornamental value, require monitoring multiple growth parameters to balance yield and aesthetics. While deep learning has been widely applied in phenotype monitoring, most studies focus on individual growth parameters, overlooking intrinsic relationships. To simultaneously monitor multiple growth parameters across the entire growth stage and different cultivars, this study develops a multi-modal multi-task phenotype monitoring network for dwarf tomatoes (TomPhenoNet). The network model utilizes top-view RGB-D images to evaluate four key growth parameters: height, leaf area, fresh weight, and the number of red fruits. TomPhenoNet generates mask images, fruit detection features, and the number of detected fruits based on RGB images. By fusing RGB-D images, mask images, and fruit detection features, and introducing the cross-stitch network, the network predicts plant height, leaf area, and fresh weight. The predicted values are further used to generate the dynamic occlusion coefficient, adjusting the number of detected fruits to accurately predict the number of red fruits. Results reveal that TomPhenoNet achieves high prediction performances, with R² values of 0.828, 0.930, 0.945, and 0.881 for plant height, leaf area, fresh weight, and the number of red fruits, respectively. Ablation experiments show that the cross-stitch network and fruit detection features improve the prediction performances of growth parameters, with TomPhenoNet combining both modules performing best. Feature importance analysis indicates the network model captures plant growth characteristics and corrects the impact of leaf occlusion from the top view. This study promotes accurate tomato monitoring and provides data support for optimizing cultivation strategies.
Why it matches plant phenotyping methodsトマトの複数形質をRGB-D画像から推定するマルチモーダル・マルチタスク手法を開発し、性能評価とアブレーション実験まで行っており、表現型取得・推定法が研究の中心である。
abstractthis study develops a multi-modal multi-task phenotype monitoring network for dwarf tomatoes (TomPhenoNet).
Precisely recognizing diseases in tomato leaves plays a crucial role in enhancing the health, productivity, and quality of tomato crops. However, disease identification methods that rely on single-mode information often face the problems of insufficient accuracy and weak generalization ability. Therefore, this paper proposes a tomato leaf disease recognition framework FCMNet based on multimodal fusion, which combines tomato leaf disease image and text description to enhance the ability to capture disease characteristics. In this paper, the Fourier-guided Attention Mechanism (FGAM) is designed, which systematically embeds the Fourier frequency-domain information into the spatial-channel attention structure for the first time, enhances the stability and noise resistance of feature expression through spectral transform, and realizes more accurate lesion location by means of multi-scale fusion of local and global features. In order to realize the deep semantic interaction between image and text modality, a Cross Vision-Language Alignment module (CVLA) is further proposed. This module generates visual representations compatible with Bert embeddings by utilizing block segmentation and feature mapping techniques. Additionally, it incorporates a probability-based weighting mechanism to achieve enhanced multimodal fusion, significantly strengthening the model's comprehension of semantic relationships across different modalities. Furthermore, to enhance both training efficiency and parameter optimization capabilities of the model, we introduce a Multi-strategy Improved Coati Optimization Algorithm (MSCOA). This algorithm integrates Good Point Set initialization with a Golden Sine search strategy, thereby boosting global exploration, accelerating convergence, and effectively preventing entrapment in local optima. Consequently, it exhibits robust adaptability and stable performance within high-dimensional search spaces. The experimental results show that the FCMNet model has increased the accuracy and precision by 2.61% and 2.85%, respectively, compared with the baseline model on the self-built dataset of tomato leaf diseases, and the recall and F1 score have increased by 3.03% and 3.06%, respectively, which is significantly superior to the existing methods. This research provides a new solution for the identification of tomato leaf diseases and has broad potential for agricultural applications.
Why it matches plant phenotyping methodsトマト葉の病斑・病害状態を画像から推定するマルチモーダル認識フレームワークを開発し、ベースライン比較で技術性能を評価しているため、植物表現型取得・推定手法が中心である。
abstractthis paper proposes a tomato leaf disease recognition framework FCMNet based on multimodal fusion
Delayed identification of crop diseases, which significantly impact agricultural yields, remains a critical challenge. Crop diseases are a major factor contributing to reducing productivity. Since leaves are the mirrors of crop health, by investigating the leaves, a prediction of crop health can be made. This study aims to predict crop disease in the vegetative growth phase with greater efficiency. The two most prominent features, color and texture of the leaves, are extracted with different techniques, followed by fuzzification of these features. Two machine learning models, the bootstrap model and the multi-class support vector machine (MSVM), are employed for disease prediction. The findings show that for multi-class disease prediction, the bootstrap model with histogram and modified co-occurrence matrix features obtains a superior average accuracy of 98.07%, while the MSVM with fuzzy features delivers an average accuracy of 80.11% in the potato crop with early blight disease.
Why it matches plant phenotyping methods葉の色・テクスチャを画像特徴として抽出し、植物病害状態を推定する手法が研究の中心であり、病害フェノタイピング手法として適格です。
abstractThe two most prominent features, color and texture of the leaves, are extracted with different techniques, followed by fuzzification of these features.
Anthesis prediction is crucial for breeding wheat. While current tools provide estimates of average anthesis at the field scale, they fail to address the needs of breeders who require accurate predictions for individual plants. Hybrid breeders have to finalize their plans for pollination at least 10 days before such flowering is due and biotechnology field trials in the United States and Australia must report to regulators 7-14 days before the first plant flowers. Currently, predicting anthesis of individual wheat plants is a labour-intensive, inefficient, and costly process. Individual wheat of the same cultivar within the same field may exhibit substantial variations in anthesis timing, due to significant variations in their immediate surroundings. In this study, we developed an efficient and cost-effective machine vision approach to predict anthesis of individual wheat plants. By integrating RGB imagery with in-situ meteorological data, our multimodal framework simplifies the anthesis prediction problem into binary or three-class classification tasks, aligning with breeders' requirements in individual wheat flowering prediction on the crucial days before anthesis. Furthermore, we incorporated a few-shot learning method to improve the model's adaptability across different growth environments and to address the challenge of limited training data. The model achieved an F1 score above 0.8 in all planting settings.
Why it matches plant phenotyping methods個体別コムギの開花(anthesis)時期をRGB画像と気象データから予測する機械視覚・ few-shot 学習手法の開発が中心であり、植物の発育状態を直接推定している。
abstractIn this study, we developed an efficient and cost-effective machine vision approach to predict anthesis of individual wheat plants.
Timely and accurate identification of plant diseases is critical to mitigating crop losses and enhancing yield in precision agriculture. This paper proposes AgriFusionNet, a lightweight and efficient deep learning model designed to diagnose plant diseases using multimodal data sources. The framework integrates RGB and multispectral drone imagery with IoT-based environmental sensor data (e.g., temperature, humidity, soil moisture), recorded over six months across multiple agricultural zones. Built on the EfficientNetV2-B4 backbone, AgriFusionNet incorporates Fused-MBConv blocks and Swish activation to improve gradient flow, capture fine-grained disease patterns, and reduce inference latency. The model was evaluated using a comprehensive dataset composed of real-world and benchmarked samples, showing superior performance with 94.3% classification accuracy, 28.5 ms inference time, and a 30% reduction in model parameters compared to state-of-the-art models such as Vision Transformers and InceptionV4. Extensive comparisons with both traditional machine learning and advanced deep learning methods underscore its robustness, generalization, and suitability for deployment on edge devices. Ablation studies and confusion matrix analyses further confirm its diagnostic precision, even in visually ambiguous cases. The proposed framework offers a scalable, practical solution for real-time crop health monitoring, contributing toward smart and sustainable agricultural ecosystems.
Why it matches plant phenotyping methods植物病害状態をRGB・マルチスペクトル画像とセンサーデータから推定する深層学習手法の開発・評価が中心であり、植物フェノタイピング手法に該当する。
abstractThis paper proposes AgriFusionNet, a lightweight and efficient deep learning model designed to diagnose plant diseases using multimodal data sources.
Global challenges such as climate change and population growth require improvements in crop monitoring models. To address these issues, this study advances the identification of potato crop phenological stages using satellite remote sensing, a field where cereals have been the primary focus. We introduce a methodology using Sentinel-1 (S1) and Sentinel-2 (S2) time series data to pinpoint critical phenological stages—emergence, canopy closure, flowering, senescence onset, and harvest timing—at the field scale. Our approach utilizes analysis of NDVI, fAPAR, and IRECI2 from S2, alongside VH and VV polarizations from S1, informed by domain knowledge of the spectral and morphological responses of potato crops. We propose the integration of NDVI and VH indices, NDVI_VH, to improve stage detection accuracy. Comparative analysis with ground-observed stages validated the method’s effectiveness, with NDVI proving to be one of the most informative indices, achieving RMSEs of 12 and 14 days for emergence and closure, and 17 days for the onset of senescence. The integrated NDVI_VH approach complemented NDVI, particularly in harvest and flowering stages, where VH enhanced accuracy, achieving an overall R2 value of 0.80. The study demonstrates the potential of combining SAR and optical data for post-season crop phenology analysis, providing insights that can inform the development of new methods and strategies to enhance on-season crop monitoring and yield forecasting.
Why it matches plant phenotyping methodsSentinel-1/2時系列からジャガイモの生育ステージを抽出する手法を開発し、地上観測で精度検証しており、植物フェノタイピングが中心である。
abstractWe introduce a methodology using Sentinel-1 (S1) and Sentinel-2 (S2) time series data to pinpoint critical phenological stages—emergence, canopy closure, flowering, senescence onset, and harvest timing—at the field scale.
A complete plant body consists of elements on different scales, including microscopic molecules, mesoscopic multicellular structures, and macroscopic tissues and organs, which are interconnected to form complex biological networks. The growth and development of plants involve the regulation of elements on different scales and their biological networks, which requires the coordinated operation of multiple molecules, cells, tissues, and organs. It is difficult to reveal the essence of multi-level life activities by a single method or technology. In recent years, the development of various novel imaging technologies has provided new approaches for revealing the complex life activities in plants. Using multi-modal imaging technologies to study the cross-scale network connections of plants from the microscopic, mesoscopic, and macroscopic levels is crucial for understanding the complex internal connections behind biological functions. This paper first summarizes multi-modal cross-scale imaging technologies, three-dimensional reconstruction, and image processing methods, outlines the basic framework of cross-scale network connection properties, and then summarizes the applications of multi-modal imaging technologies in elucidating plant multi-scale networks. Finally, this review systematically integrates the combined analysis of cross-scale 3D spatial structural data and single-cell omics, laying a theoretical foundation for the innovation of novel plant imaging technologies. Furthermore, it provides a new research paradigm for in-depth exploration of the interaction mechanisms among cross-scale elements and the principles of biological network connectivity in plant life activities.
Why it matches plant phenotyping methods植物のマルチモーダル画像技術、3D再構成、画像処理を体系的にレビューしており、植物の構造・状態を取得する画像ベース手法が中心です。
abstractThis paper first summarizes multi-modal cross-scale imaging technologies, three-dimensional reconstruction, and image processing methods
Pests and diseases significantly impact the growth and development of crops. When attempting to precisely identify disease characteristics in crop images through dialogue, existing multimodal models face numerous challenges, often leading to misinterpretation and incorrect feedback regarding disease information. This paper proposed a large language model for multimodal identification of crop diseases and pests, which can be called LLMI-CDP. It builds up on the VisualGLM model and introduces improvements to achieve precise identification of agricultural crop disease and pest images, along with providing professional recommendations for relevant preventive measures. The use of Low-Rank Adaptation (LoRA) technology, which adjusts the weights of pre-trained models, achieves significant performance improvements with a minimal increase in parameters. This ensures the precise capture and efficient identification of crop pest and disease characteristics, greatly enhancing the model's application flexibility and accuracy in the field of pest and disease recognition. Simultaneously, the model incorporates the Q-Former framework for effective modal alignment between language models and image features. Through this approach, the LLMI-CDP model is able to more deeply understand and process the complex relationships between language and visual information, further enhancing its performance in multimodal recognition tasks. Experiments are carried out in the homemade datasets, The results demonstrate that the LLMI-CDP model surpasses five leading multimodal large language models in relevant evaluation metrics, confirming its outstanding performance in Chinese multimodal dialogues related to agriculture.
Why it matches plant phenotyping methods作物画像から病害・害虫の状態を識別するマルチモーダル手法を開発し、自作データセットで既存モデルと比較評価しており、植物の病害状態の取得・推定が中心です。
titleA large language model for multimodal identification of crop diseases and pests.
Autonomous mobile robotic solutions are increasingly being explored in precision agriculture to aid human workers in labour-intensive or repetitive tasks. Moreover, the emergence of foundation models in vision-based AI domain presents an opportunity to perform automated interpretation of in-field collected data. This study presents a cost-effective mobile robotic research platform designed for autonomous vineyard inspection: it integrates mission planning, real-world navigation and a post-processing pipeline of multimodal data. The system, based on the Leo rover, is equipped with LiDAR, RGB cameras and GNSS-visual-inertial positioning, ensuring reliable operation in GNSS-degraded vineyard environments. We propose a novel methodology for automating several stages of the workflow using various open and in-situ collected data. The robotic platform and processing pipeline were validated through simulation and field experiments, demonstrating its capability for autonomous navigation, 3D reconstruction, AI-based fruit detection and an initial plant health assessment through Large Multimodal Models (LMM). Results show that while 3D mapping provides highresolution spatial data, AI-driven object detection and vision models require further domain adaptation for reaching reliable and trustable operation. The study highlights the feasibility of cost-effective mobile robotic solutions in vineyard monitoring and the potential of integrating AI to enhance agricultural automation.
Why it matches plant phenotyping methods自律型ロボットとマルチモーダル処理パイプラインを開発・検証し、3D再構成、果実検出、植物健全性評価という植物状態の取得を中核的に扱っているため。
abstractThis study presents a cost-effective mobile robotic research platform designed for autonomous vineyard inspection
Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate disease identification is crucial for improving crop yield and quality. While most existing deep learning-based methods focus primarily on image datasets for disease recognition, they often overlook the complementary role of textual features in enhancing visual understanding. To address this problem, we proposed a cross-modal data fusion via a vision-language model for crop disease recognition. Our approach leverages the Zhipu.ai multi-model to generate comprehensive textual descriptions of crop leaf diseases, including global description, local lesion description, and color-texture description. These descriptions are encoded into feature vectors, while an image encoder extracts image features. A cross-attention mechanism then iteratively fuses multimodal features across multiple layers, and a classification prediction module generates classification probabilities. Extensive experiments on the Soybean Disease, AI Challenge 2018, and PlantVillage datasets demonstrate that our method outperforms state-of-the-art image-only approaches with higher accuracy and fewer parameters. Specifically, with only 1.14M model parameters, our model achieves a 98.74%, 87.64% and 99.08% recognition accuracy on the three datasets, respectively. The results highlight the effectiveness of cross-modal learning in leveraging both visual and textual cues for precise and efficient disease recognition, offering a scalable solution for crop disease recognition.
Why it matches plant phenotyping methods作物葉の病徴を画像・テキストから認識するマルチモーダル手法の開発と評価が研究の中心であり、植物の病害状態を直接推定している。
abstractwe proposed a cross-modal data fusion via a vision-language model for crop disease recognition.
Reproduction assets foundThe paper's Data Availability Statement explicitly links the public image datasets used for its crop disease recognition experiments: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). These are the phenotyping image inputs directly used in this study. The AI Challenge 2018 dataset is also公开Dataset · publicThe soybean
dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3 (accessed on 1 April 2025).Open asset ↗Dryad · 10.5061/dryad.41ns1rnj3pdf-page:12 lines:1-58Dataset · publicplantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-Open asset ↗Kagglepdf-page:12 lines:1-58Plant phenotyping relevance match · UnverifiedEurope PMC · Crossref · checked 6 Sept 2026
Abstract Potato crops are a vital part of global food security. Potato leaf and crop health play a crucial role in determining the yield and quality of potato production. This paper presents a novel approach to potato disease detection by integrating a Vision Transformer (ViT) model with a Large Language Model (LLM) for enhanced classification of potato plant diseases. We developed a multi-modal pipeline that not only accurately identifies diseases affecting potato leaves and tubers but also provides contextual explanations for the diagnoses. Experimental results demonstrate that our integrated approach outperforms traditional individual models, with the potato leaf disease classifier achieving 99.44% validation accuracy and the potato tuber disease classifier reaching 76.19% accuracy when trained separately, while the combined model maintains excellent performance of 95.06% on the validation set. The fusion of computer vision with Mistral AI's LLM capabilities creates an interpretable system that can assist agricultural experts with both disease identification and recommended treatment actions. This paper contributes to the growing field of AI-assisted agriculture by demonstrating how multi-modal deep learning systems can provide more comprehensive solutions to potato disease management challenges, potentially reducing crop losses and improving food security.
Why it matches plant phenotyping methodsジャガイモ葉・塊茎の病徴を画像から分類するマルチモーダル手法が研究の中心であり、植物の病害状態を直接推定するため、植物フェノタイピング手法として適格です。
abstractThis paper presents a novel approach to potato disease detection by integrating a Vision Transformer (ViT) model with a Large Language Model (LLM) for enhanced classification of potato plant diseases.
Addressing the global malnutrition crisis requires precise and timely diagnostics of plant stresses to enhance the quality and yield of nutrient-rich crops, such as tomatoes. Soft wearable sensors offer a promising approach by continuously monitoring plant physiology. However, challenges remain in identifying direct physiological indicators of plant stresses, hindering the development of accurate diagnostic models for predicting symptom progression. Here, we introduce a machine-learning-powered spectral-dominant multimodal soft wearable system (MapS-Wear) for precise, long-term, and early-stage diagnosis of stresses in tomatoes. MapS-Wear continuously tracks leaf surrounding temperature, humidity, and unique in-situ transmission spectra, which are critical stress-related indicators. The machine learning framework processes these multimodal data to predict gradual stress progression and diagnose nutrient deficiencies in plants over 10 days earlier than conventional computer vision methods. Moreover, MapS-Wears enables portable and large-scale screening of grafted tomato varieties in greenhouses, accelerating the identification of compatible grafting combinations. This demonstration highlights the potential for high-throughput plant phenotyping and yield improvement.
Why it matches plant phenotyping methods植物ストレスの生理状態を連続センシングし、機械学習で早期診断・進行予測するウェアラブル計測システムが研究の中心であり、植物フェノタイピング手法として明確に該当する。
abstractHere, we introduce a machine-learning-powered spectral-dominant multimodal soft wearable system (MapS-Wear) for precise, long-term, and early-stage diagnosis of stresses in tomatoes.
Reproduction assets foundThe paper's Data and materials availability statement explicitly deposits the tomato leaf photos, transmission spectral data, and ML algorithms on Zenodo, matching an allowed URL.Dataset · publicThe photos of tomato leaves in different health statuses, the transmission spectral data of these leaves, and the ML algorithms are openly available on Zenodo ( https://zenodo.org/doi/10.5281/zenodo.15192884 ).Open asset ↗Zenodo · 10.5281/zenodo.15192884lines:129-274Plant phenotyping relevance match · UnverifiedCrossref · checked 14 Sept 2026
Published27 Jun 2025Informatyka, Automatyka, Pomiary w Gospodarce i Ochronie ŚrodowiskaCited by 3 · OpenAlex ↗
Advancements in genomics and artificial intelligence are transforming precision agriculture by enabling early stress detection and adaptive crop management. Integrating genomic analysis, image-based stress detection, and real-time environmental monitoring, this approach assesses plant responses to stress factors such as drought and disease. A BERT-based model processes genomic data, while computer vision identifies visual stress indicators like wilting and discoloration. IoT sensors track environmental parameters such as soil moisture, temperature, and humidity, refining predictions and optimizing intervention strategies. The system leverages multimodal data fusion to enhance decision-making, improving the accuracy of stress detection and mitigation strategies. Machine learning models continuously adapt by learning from historical and real-time data, making recommendations more precise over time. A web-based platform allows users to upload plant images and environmental data for real-time analysis, generating personalized recommendations for irrigation, fertilization, and disease management. The platform's intuitive interface ensures accessibility for farmers and agricultural experts, facilitating widespread adoption. By combining AI, genomics, and IoT, this system enhances crop health, maximizes yield, and promotes sustainable farming through proactive, data-driven decision-making. Ultimately, it aims to reduce resource waste, mitigate crop losses, and support scalable, technology-driven agricultural solutions.
Why it matches plant phenotyping methods画像から萎れや変色などの植物ストレス状態を検出するAI手法と、画像・ゲノム・環境センサの統合プラットフォームが中心であり、植物状態の推定を技術的に扱っているため。
abstractIntegrating genomic analysis, image-based stress detection, and real-time environmental monitoring, this approach assesses plant responses to stress factors such as drought and disease.
With the advancement of Agriculture 4.0 and the ongoing transition toward sustainable and intelligent agricultural systems, deep learning-based multimodal fusion technologies have emerged as a driving force for crop monitoring, plant management, and resource conservation. This article systematically reviews research progress from three perspectives: technical frameworks, application scenarios, and sustainability-driven challenges. At the technical framework level, it outlines an integrated system encompassing data acquisition, feature fusion, and decision optimization, thereby covering the full pipeline of perception, analysis, and decision making essential for sustainable practices. Regarding application scenarios, it focuses on three major tasks—disease diagnosis, maturity and yield prediction, and weed identification—evaluating how deep learning-driven multisource data integration enhances precision and efficiency in sustainable farming operations. It further discusses the efficient translation of detection outcomes into eco-friendly field practices through agricultural navigation systems, harvesting and plant protection robots, and intelligent resource management strategies based on feedback-driven monitoring. In addressing challenges and future directions, the article highlights key bottlenecks such as data heterogeneity, real-time processing limitations, and insufficient model generalization, and proposes potential solutions including cross-modal generative models and federated learning to support more resilient, sustainable agricultural systems. This work offers a comprehensive three-dimensional analysis across technology, application, and sustainability challenges, providing theoretical insights and practical guidance for the intelligent and sustainable transformation of modern agriculture through multimodal fusion.
Why it matches plant phenotyping methods植物の病害診断、成熟度・収量予測を含むマルチモーダルな観測・融合手法を体系的にレビューしており、植物状態・形質の取得と解析方法が中心である。
abstractThis article systematically reviews research progress from three perspectives: technical frameworks, application scenarios, and sustainability-driven challenges.
Field / plotMultimodalMultispectral / hyperspectralWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenology
Crop phenology describes the physiological development stages of crops from planting to harvest which is valuable information for decision makers to plan and adapt agricultural management strategies. In the era of big Earth observation data ubiquity, attempts have been made to accurately detect crop phenology using Remote Sensing (RS) and high resolution weather data. However, most studies have focused on large scale predictions of phenology or developed methods which are not adequate to help crop modeler communities on leveraging Sentinel-1 and Sentinal-2 data and fusing them with high resolution climate data, using a novel framework. For this, we trained a Machine Learning (ML) LightGBM model to predict 13 phenological stages for eight major crops across Germany at 20 m scale. Observed phenologies were taken from German national phenology network (German Meteorological Service; DWD) between 2017 and 2021. We proposed a thorough feature selection analysis to find the best combination of RS and climate data to detect phenological stages. At national scale, predicted phenology resulted in a reasonable precision of R 2 > 0.43 and a low Mean Absolute Error of 6 days, averaged over all phenological stages and crops. The spatio-temporal analysis of the model predictions demonstrates its transferability across different spatial and temporal context of Germany. The results indicated that combining radar sensors with climate data yields a very promising performance for a multitude of practical applications. Moreover, these improvements are expected to be useful to generate highly valuable input for crop model calibrations and evaluations, facilitate informed agricultural decisions, and contribute to sustainable food production to address the increasing global food demand. • A ML framework is proposed to detect phenological stages of various crops. • Proposed ML model can detect crop phenology for eight crops in Germany. • Fusing radar sensor with climate data improved phenology identification. • Geospatial and climate parameters plays major role in phenology detection. • Fearture optimization can help to find the most important features in ML modeling.
Why it matches plant phenotyping methodsSentinel-1/2と気候データを融合し、機械学習で作物の13段階の生育フェノロジーを推定する枠組みが研究の中心であり、植物状態の取得・推定方法を扱っている。
abstractWe proposed a thorough feature selection analysis to find the best combination of RS and climate data to detect phenological stages.
Accurate and timely crop disease detection is critical for reducing agricultural losses and ensuring food security in low-resource settings. Traditional diagnostic methods, such as manual inspections, are often inefficient and error-prone. Existing deep learning models (e.g., ResNet50, Inception V3) struggle with computational inefficiency and poor generalizability in real-world farming contexts. This study proposes a lightweight multimodal fusion model integrating EfficientNetV2 and MobileNetV2, optimized for edge deployment. The architecture leverages compound scaling and feature fusion to recognize subtle disease patterns, and it was fine-tuned on a globally diverse dataset (PlantVillage and field-collected leaf images). The proposed model achieved state-leading metrics (99.0% accuracy, 0.993 precision, 0.990 F1-score, AUC = 0.999997), outperforming benchmarks like ShuffleNet and DenseNet50 (ranked 2nd–6th). Statistical validation via the Kruskal-Wallis test confirmed significant performance differences across models (H=614.90, p=1.4237e−129), with Bayesian analysis showing a 100% superiority probability over DenseNet50. Notably, the model exhibited the lowest confidence variance (0.000012) compared to alternatives (0.000014–0.000032), demonstrating unmatched prediction stability. Deployment on low-end mobile devices posed challenges such as computational constraints and offline usability. However, the TensorFlow Lite-powered mobile app addressed these limitations, offering real-time, offline disease classification with 0.094-second inference latency on devices with ≤2GB RAM. Validated on 249 unseen field images (95.98% accuracy), this solution bridges the gap between high-performance deep learning and real-world agricultural needs, empowering smallholder farmers with an accessible and scalable tool.
Why it matches plant phenotyping methods葉画像から植物病害を推定する深層学習モデルとモバイル実装を開発し、ベンチマーク比較および未見フィールド画像で検証しており、植物表現型(病害状態)の取得・推定が中心である。
abstractThis study proposes a lightweight multimodal fusion model integrating EfficientNetV2 and MobileNetV2, optimized for edge deployment.
Deep learning has transformed computer vision for precision agriculture, yet apple orchard monitoring remains limited by dataset constraints. The lack of diverse, realistic datasets and the difficulty of annotating dense, heterogeneous scenes. Existing datasets overlook different growth stages and stereo imagery, both essential for realistic 3D modeling of orchards and tasks like fruit localization, yield estimation, and structural analysis. To address these gaps, we present AppleGrowthVision, a large-scale dataset comprising two subsets. The first includes 9,317 high resolution stereo images collected from a farm in Brandenburg (Germany), covering six agriculturally validated growth stages over a full growth cycle. The second subset consists of 1,125 densely annotated images from the same farm in Brandenburg and one in Pillnitz (Germany), containing a total of 31,084 apple labels. AppleGrowthVision provides stereo-image data with agriculturally validated growth stages, enabling precise phenological analysis and 3D reconstructions. Extending MinneApple with our data improves YOLOv8 performance by 7.69 % in terms of F1-score, while adding it to MinneApple and MAD boosts Faster R-CNN F1-score by 31.06 %. Additionally, six BBCH stages were predicted with over 95 % accuracy using VGG16, ResNet152, DenseNet201, and MobileNetv2. AppleGrowthVision bridges the gap between agricultural science and computer vision, by enabling the development of robust models for fruit detection, growth modeling, and 3D analysis in precision agriculture. Future work includes improving annotation, enhancing 3D reconstruction, and extending multimodal analysis across all growth stages.
Why it matches plant phenotyping methodsリンゴの生育段階・果実・樹体構造を対象とする大規模ステレオ画像データセットを構築し、果実検出、フェノロジー分析、3D再構成モデルの評価に用いており、表現型取得・解析基盤が中心である。
titleAppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards
Plant diseases and abiotic stressors threaten global food security, necessitating rapid, noninvasive detection methods. Current single-modality imaging techniques, such as RGB (spectral) or Optical Coherence Tomography (OCT; structural), face limitations in misclassification and incomplete feature extraction. This study proposes a multimodal fusion approach integrating OCT and RGB imaging to enhance machine learning (ML)-based classification of rice leaf health. OCT provided high-resolution cross-sectional microstructural data (axial resolution: ∼4 µm, depth range: ∼3.2 mm), revealing statistically significant differences in optical attenuation coefficients (healthy: 3.65 ± 0.36 mm−1, dry: 4.75 ± 0.4 mm−1, diseased: 5.49 ± 0.52 mm−1; p < 0.0001) and leaf thickness (healthy: 437.4 ± 38.32 µm, dry: 331.4 ± 21.83 µm, diseased: 266 ± 27.91 µm). RGB imaging captured spectral reflectance variations across visible wavelengths (480–650 nm), achieving 96.7% classification accuracy. Neural networks trained on combined OCT-RGB data achieved 99.6% accuracy, outperforming single-modality models (OCT: 88.9%, RGB: 96.7%). The fusion strategy leveraged complementary features: OCT quantified subsurface structural degradation, while RGB detected surface-level biochemical changes. This dual-modality approach demonstrates superior diagnostic precision, offering a scalable, noninvasive solution for early disease detection and stress phenotyping in precision agriculture.
Why it matches plant phenotyping methodsOCTとRGB画像を融合してイネ葉の病害・ストレス状態を定量・分類する画像ベース手法を開発し、単一モダリティとの精度比較で検証しているため、植物表現型取得が研究の中心である。
abstractThis study proposes a multimodal fusion approach integrating OCT and RGB imaging to enhance machine learning (ML)-based classification of rice leaf health.
Accurate measurement of structural phenotypes, such as plant height and canopy width, is crucial for the scientific management of lettuce cultivation in Plant Factories with Artificial Lighting (PFALs). In this study, we developed a multimodal image fusion model using visible images (RGB), depth images (Depth), and infrared images (IR) to extract lettuce phenotypes in PFAL environments. We proposed a Residual Space information enhancement module (DRS) and a fusion feature supplement method with adaptive weight optimization for IR features (IRC) to address the weak spatial perception of traditional RGB-based models and the feature loss of RGB due to illumination disturbance. Three lettuce varieties (Bixiao, Huqian, and Mondai) were selected as experimental subjects to evaluate the robustness of our proposed model. In ablation experiments, the benchmark model improved by DRS increased by 1.6% and 0.9% in terms of MAP0.75 and MAP0.5:0.95, respectively. The benchmark model improved by IRC increased by 0.2%, 0.6%, and 1.2% in terms of MAP0.5, MAP0.75, and MAP0.5:0.95, respectively. Furthermore, MAP0.5, MAP0.75, and MAP0.5:0.95 values increased by 0.3%, 3.3%, and 2.3% when the two modules were combined, respectively. Compared with manually measured plant height and canopy width, the Root Mean Square Error (RMSE) of the average plant height prediction results for the three varieties is 0.74, and the Mean Squared Error (MSE) is 0.55. For canopy width, the RMSE of the model’s prediction results was 0.70, and the MSE was 0.49. In the lighting influence experiment, our method outperformed the unimproved model by approximately 0.3–4% in terms of MAP0.5, MAP0.75, and MAP0.5:0.95 across multiple datasets. Our proposed model effectively addresses lighting disturbance, enhances the robustness of the baseline model against varying lighting conditions, improves spatial perception capability, facilitates the separation of adjacent plant features in the model’s feature extraction stage, enhances the model’s detection ability, and ultimately improves phenotype extraction capability.
Why it matches plant phenotyping methodsマルチモーダル画像融合モデルを開発し、レタスの草丈・群落幅という植物表現型を画像から抽出・推定しており、手法開発と技術評価が研究の中心である。
abstractIn this study, we developed a multimodal image fusion model using visible images (RGB), depth images (Depth), and infrared images (IR) to extract lettuce phenotypes in PFAL environments.
The effective tiller number of wheat (ETNW) is one of the three main factors affecting wheat yield. Most traditional methods for counting tillers are manual, which is inefficient and challenging to implement on a wide scale. Existing remote sensing techniques for tiller counting are typically focused on individual plants or small plot areas, often utilizing high-resolution sensors for close-range monitoring. This approach limits the scalability and applicability for large-scale field environments. Moreover, most studies on wheat tiller monitoring have concentrated on specific crop varieties or single-variable conditions, with little research conducted on estimating or monitoring ETNW in large-scale field scenarios using consumer-grade unmanned aerial vehicle (UAV) platforms. To address these issues, we propose a multi-modal fusion-driven machine learning method to enhance the performance of wheat tillering monitoring. This study was conducted using two experimental setups to ensure robust and comprehensive data collection. Experiment 1 (Exp.1) focused on water and nitrogen coupling conditions, while Experiment 2 (Exp.2) included both nitrogen-deficient and nitrogen-sufficient treatments. Each experiment involved multiple wheat varieties to account for genotypic variability. UAV data, including multispectral and RGB imagery, was collected across different growth stages under varying irrigation and nitrogen conditions to ensure the generalizability of the proposed model. This method employs spectral correlation analysis (SCA) to select features strongly correlated with the number of tillers and subsequently fuses these features to construct a multi-modal machine learning model based on vegetation index (VI), color index (CI), multispectral texture features (TF1), and RGB texture features (TF2) derived from UAV data. Three machine learning models are applied: Random Forest Regression (RFR), Partial Least Squares Regression (PLSR), and Support Vector Regression (SVR). The results demonstrated that the vegetation index-driven machine learning model (VI-RFR) achieved the best performance, with an R² of 0.85, RMSE of 348.60, and NRMSE of 0.367, followed by the color index model (CI-RFR, R² = 0.83, RMSE = 313.18, NRMSE = 0.330), the multispectral texture feature model (TF1-RFR, R² = 0.81, RMSE = 357.12, NRMSE = 0.376), and the RGB texture feature model (TF2-RFR, R² = 0.80, RMSE = 373.69, NRMSE = 0.393). Compared to using a single data source, data fusion significantly improved model accuracy, particularly when complementary data sources were combined. Specifically, the VI&CI combination achieved an R² improvement of 6.4 %-19.1 %, an RMSE reduction of 2.8 %-9%, and an NRMSE decrease of 1.56 %-8.24 %. The VI&TF1 combination exhibited an R² increase of 7.8 %–32.6 %, an RMSE reduction of 8.5 %-28.8 %, and an NRMSE decrease of 2.20 %-8.24 %. The VI&TF2 combination showed an R² increase of 3.8 %-40.5 %, an RMSE reduction of 2.8 %-27.6 %, and an NRMSE decrease of 2.68 %-21.61 %. The CI&TF1 combination resulted in an R² improvement of 5.2 %-25.6 %, an RMSE reduction of 7.4 %-24.1 %, and an NRMSE decrease of 2.04 %-19.41 %. The CI&TF2 combination achieved an R² increase of 12.5 %–33.3 %, an RMSE reduction of 5.4 %–23.4 %, and an NRMSE decrease of 1.17 %-18.94 %. The TF1&TF2 combination achieved an R² improvement of 13.9 %–33.3 %, an RMSE reduction of 10.9 %-27.6 %, and an NRMSE decrease of 5.01 %-21.61 %. However, increasing the number of data sources does not necessarily lead to higher model accuracy. The VI&CI&TF1 fusion model (RFR, R² = 0.90, RMSE = 288.39, NRMSE = 0.304) demonstrated the best performance, surpassing the combination of all four features. The machine learning models exhibited systematic variations, with the performance ranking as follows: RFR > SVR > PLSR, regardless of the data fusion strategy employed. Moreover, the best model (VI&CI&TF1-RFR) demonstrated adaptability across different growth stages, irrigation conditions, and nitrogen fertilizer applications. Notably, the model maintained robustness even in extreme environments with complete nitrogen deficiency. The findings of this study provide technical support for crop phenotyping, variety selection, and precision agriculture management.
Why it matches plant phenotyping methodsUAVのマルチスペクトル・RGB画像から小麦の有効分げつ数を推定する特徴融合・機械学習手法を開発し、複数条件・品種で性能評価しており、植物表現型の取得・推定が研究の中心である。
abstractTo address these issues, we propose a multi-modal fusion-driven machine learning method to enhance the performance of wheat tillering monitoring.
Plant phenotyping relevance match · UnverifiedOpenAlex · Crossref · Europe PMC · checked 6 Sept 2026
Introduction: Detecting plant stress is a critical challenge in agriculture, where early intervention is essential to enhance crop resilience and maximize yield. Conventional single-mode approaches often fail to capture the complex interplay of plant health stressors. Methods: This review integrates findings from recent advancements in Multi-Mode Analytics (MMA), which employs spectral imaging, image-based phenotyping, and adaptive computational techniques. It integrates machine learning, data fusion, and hyperspectral technologies to improve analytical accuracy and efficiency. Results: MMA approaches have shown substantial improvements in the accuracy and reliability of early interventions. They outperform traditional methods by effectively capturing complex interactions among various abiotic stressors. Recent research highlights the benefits of MMA in enhancing predictive capabilities, which facilitates the development of timely and effective intervention strategies to boost agricultural productivity. Discussion: The advantages of MMA over conventional single-mode techniques are significant, particularly in the detection and management of plant stress in challenging environments. Integrating advanced analytical methods supports precision agriculture by enabling proactive responses to stress conditions. These innovations are pivotal for enhancing food security in terrestrial and space agriculture, ensuring sustainability and resilience in food production systems.
Why it matches plant phenotyping methods植物ストレス評価のためのスペクトル画像、画像ベース表現型解析、機械学習・データ融合を中心に扱うレビューであり、植物表現型計測手法のレビューとして中心的です。
titleA systematic review of multi-mode analytics for enhanced plant stress evaluation
To address the low estimation accuracy of deep learning-based crop yield image recognition methods under untrained shooting distances, this study proposes a shooting distance adaptive crop yield estimation method by fusing RGB and depth image information through multi-modal data fusion. Taking strawberry fruit fresh weight as an example, RGB and depth image data of 348 strawberries were collected at nine heights ranging from 70 to 115 cm. First, based on RGB images and shooting height information, a single-modal crop yield estimation model was developed by training a convolutional neural network (CNN) after cropping strawberry fruit images using the relative area conversion method. Second, the height information was expanded into a data matrix matching the RGB image dimensions, and multi-modal fusion models were investigated through input-layer and output-layer fusion strategies. Finally, two additional approaches were explored: direct fusion of RGB and depth images, and extraction of average shooting height from depth images for estimation. The models were tested at two untrained heights (80 cm and 100 cm). Results showed that when using only RGB images and height information, the relative area conversion method achieved the highest accuracy, with R2 values of 0.9212 and 0.9304, normalized root mean square error (NRMSE) of 0.0866 and 0.0814, and mean absolute percentage error (MAPE) of 0.0696 and 0.0660 at the two untrained heights. By further incorporating depth data, the highest accuracy was achieved through input-layer fusion of RGB images with extracted average height from depth images, improving R2 to 0.9475 and 0.9384, reducing NRMSE to 0.0707 and 0.0766, and lowering MAPE to 0.0591 and 0.0610. Validation using a developed shooting distance adaptive crop yield estimation platform at two random heights yielded MAPE values of 0.0813 and 0.0593. This model enables adaptive crop yield estimation across varying shooting distances, significantly enhancing accuracy under untrained conditions.
Why it matches plant phenotyping methodsRGB・深度画像とCNNを用いてイチゴ果実重量(収量)を推定する撮影距離適応型の方法とプラットフォームを開発・検証しており、表現型取得・推定が研究の中心である。
abstractthis study proposes a shooting distance adaptive crop yield estimation method by fusing RGB and depth image information through multi-modal data fusion.
Early diagnosis, the correct diagnosis of plant diseases is important to ensure sustainable agriculture and the minimalization of the loss of production. Traditional approaches of plant disease detection, which involve manual inspection and single modal imaging, are highly cumbersome, erroneous and lack in capturing the niche characteristics of the disease. Some recent achievements of deep learning advocate for possible automatic plant disease diagnosis; however, still most of the current models are plagued from low generalization capability, high computational cost and the issue of real time implementation. To alleviate these difficulties, this article introduces a brand-new multiple-mode deep learning framework, that combines RGB, hyperspectral and thermal imaging to take on the task of setting up precision and efficiency for plant disease detection. The described framework makes use of EfficientNet-based CNN for spatial feature extraction from RGB images, 1D-CNN for hyperspectral spectral feature learning and Vision Transformers (ViT) for learning long-range contextual dependencies. Above sensor- features are fused by Means of weighted summation methodology, dynamically adjusts contribution of per modality to Obtain endurance and accurate. To achieve real-time performance, the model is optimized via quantization, knowledge distillation and model pruning, with a substantial decrease in its computational load. The final optimal model is implemented in NVIDIA Jetson Nano to allow low-latency inference supporting high precision agriculture. The results of the experimental results show, the proposed multi-modal framework has achieved 97.8% accuracy, 96.5% precision, 95.7% recall and 96.1% score of F, all far exceed traditional deep learning models of ResNet-50, VGG-16, EfficientNet and Vision Transformers (ViT). Moreover, the framework offers inferences in 20 milliseconds, which makes it really suitable for real-time applications. Accomplishing a successful integration of multi-modal data fusion and model optimization not only increase classification performance, but also makes the solution/matter practical and deployable in real-world agricultural environment. The proposed framework provides a hopeful solution to smart farming, which provides a possibility of detecting disease early and managing effectively the crops.
Why it matches plant phenotyping methodsRGB・ハイパースペクトル・熱画像から植物病害を推定するマルチモーダル手法を開発し、精度・推論速度を評価しているため、植物フェノタイピング手法が中心である。
abstractthis article introduces a brand-new multiple-mode deep learning framework, that combines RGB, hyperspectral and thermal imaging to take on the task of setting up precision and efficiency for plant disease detection.
Plant height and SPAD values are critical indicators for evaluating peanut morphological development, photosynthetic efficiency, and yield optimization. Recent unmanned aerial vehicle (UAV) technology advancements have enabled high-throughput phenotyping at field scales. As a globally strategic oilseed crop, peanut plays a vital role in ensuring food and edible oil security. This study aimed to develop an optimized estimation framework for peanut plant height and SPAD values through machine learning-driven integration of UAV multi-source data while evaluating model generalizability across temporal and spatial domains. Multispectral UAV and ground data were collected across four growth stages (2023–2024). Using spectral indices and Texture features, four models (PLSR, SVM, ANN, RFR) were trained on 2024 data and independently validated with 2023 datasets. The ensemble machine learning models (RFR) significantly enhanced estimation accuracy (R2 improvement: 3.1–34.5%) and robustness compared to the linear model (PLSR). Feature stability analysis revealed that combined spectral-textural features outperformed single-feature approaches. The SVM model achieved superior plant height prediction (R2 = 0.912, RMSE = 2.14 cm), while RFR optimally estimated SPAD values (R2 = 0.530, RMSE = 3.87) across heterogeneous field conditions. This UAV-based multi-modal integration framework demonstrates significant potential for temporal monitoring of peanut growth dynamics.
Why it matches plant phenotyping methodsUAVマルチソースデータと機械学習により、ピーナッツの草丈およびSPAD値を推定するフレームワークを開発・検証しており、表現型取得・抽出手法が研究の中心である。
abstractThis study aimed to develop an optimized estimation framework for peanut plant height and SPAD values through machine learning-driven integration of UAV multi-source data while evaluating model generalizability across temporal and spatial domains.
Strawberry grading by picking robots can eliminate the manual classification, reducing labor costs and minimizing the damage to the fruit. Strawberry size or weight is a key factor in grading, with accurate weight estimation being crucial for proper classification. In this paper, we collected 1521 sets of strawberry RGB-D images using a depth camera and manually measured the weight and size of the strawberries to construct a training dataset for the strawberry weight regression model. To address the issue of incomplete depth images caused by environmental interference with depth cameras, this study proposes a multimodal point cloud completion method specifically designed for symmetrical objects, leveraging RGB images to guide the completion of depth images in the same scene. The method follows a process of locating strawberry pixel regions, calculating centroid coordinates, determining the symmetry axis via PCA, and completing the depth image. Based on this approach, a multimodal fusion regression model for strawberry weight estimation, named MMF-Net, is developed. The model uses the completed point cloud and RGB image as inputs, and extracts features from the RGB image and point cloud by EfficientNet and PointNet, respectively. These features are then integrated at the feature level through gradient blending, realizing the combination of the strengths of both modalities. Using the Percent Correct Weight (PCW) metric as the evaluation standard, this study compares the performance of four traditional machine learning methods, Support Vector Regression (SVR), Multilayer Perceptron (MLP), Linear Regression, and Random Forest Regression, with four point cloud-based deep learning models, PointNet, PointNet++, PointMLP, and Point Cloud Transformer, as well as an image-based deep learning model, EfficientNet and ResNet, on single-modal datasets. The results indicate that among traditional machine learning methods, the SVR model achieved the best performance with an accuracy of 77.7% (PCW@0.2). Among deep learning methods, the image-based EfficientNet model obtained the highest accuracy, reaching 85% (PCW@0.2), while the PointNet + + model demonstrated the best performance among point cloud-based models, with an accuracy of 54.3% (PCW@0.2). The proposed multimodal fusion model, MMF-Net, achieved an accuracy of 87.66% (PCW@0.2), significantly outperforming both traditional machine learning methods and single-modal deep learning models in terms of precision.
Why it matches plant phenotyping methodsイチゴのRGB-D画像から重量・サイズを推定する画像/点群解析手法を開発し、複数モデルと比較評価しており、植物形質取得が中心である。
abstractthis study proposes a multimodal point cloud completion method specifically designed for symmetrical objects
Rapid and accurate plant phenotyping is vital to plant breeding and monitoring. Hyperspectral imaging (HSI) is the popular phenotypic technique to acquire spectral and spatial information of plants. However, the close-range HSI of plant canopies is greatly affected by the complex interaction of canopy geometry with illumination, which leads to biased or contaminated spectral information. Thus, the mitigation of these effects is imperative but challenging. In this study, a three-dimensional (3D) spectral compensation method on close-range canopy HSI was proposed, to correct the reflectance of canopies affected by imaging distance and leaf angle. First, the hyperspectral and depth images of canopies were registered and fused to generate hyperspectral 3D point clouds. Next, the full-spectrum reflectance on canopies was compensated based on the depth and angle information provided by the hyperspectral 3D point clouds. Then, the performance of spectral compensation results was evaluated by cluster analysis, spectral curve validation, and chlorophyll regression. Results on two plant types (perilla and tea seedlings) showed that after spectral compensation, the reflectance variations within canopies reduced greatly, the reflectance of whole canopies became more homogeneous, with a dominant cluster accounting for over 67% pixels of canopies. And using the mean reflectance curves of in vitro flattened leaves as the reference, the canopy reflectance after compensation were closer to the reference level, that the Euclidean Distance (ED) between them reduced by 50.6%. The determination coefficient (R²) for chlorophyll regression after compensation reached 0.75, increasing about 17% compared to that before compensation. The overall results demonstrated that the proposed 3D spectral compensation method was effective in mitigating the effects of imaging distance and leaf angle on plant canopies in close-range HSI. This could further facilitate the revelation of plant optical characteristics, which is of high significance for the accurate close-range plant phenotyping.
Why it matches plant phenotyping methods植物キャノピーの近接ハイパースペクトル画像に対する3Dスペクトル補償法を開発し、複数の評価で性能検証しており、表現型取得・抽出手法が研究の中心である。
abstractIn this study, a three-dimensional (3D) spectral compensation method on close-range canopy HSI was proposed, to correct the reflectance of canopies affected by imaging distance and leaf angle.
Indoor gardening within sustainable buildings offers a transformative solution to urban food security and environmental sustainability. By 2030, urban farming, including Controlled Environment Agriculture (CEA) and vertical farming, is expected to grow at a compound annual growth rate (CAGR) of 13.2% from 2024 to 2030, according to market reports. This growth is fueled by advancements in Internet of Things (IoT) technologies, sustainable innovations such as smart growing systems, and the rising interest in green interior design. This paper presents a novel framework that integrates computer vision, machine learning (ML), and environmental sensing for the automated monitoring of plant health and growth. Unlike previous approaches, this framework combines RGB imagery, plant phenotyping data, and environmental factors such as temperature and humidity, to predict plant water stress in a controlled growth environment. The system utilizes high-resolution cameras to extract phenotypic features, such as RGB, plant area, height, and width while employing the Lag-Llama time series model to analyze and predict water stress. Experimental results demonstrate that integrating RGB, size ratios, and environmental data significantly enhances predictive accuracy, with the Fine-tuned model achieving the lowest errors (MSE = 0.420777, MAE = 0.595428) and reduced uncertainty. These findings highlight the potential of multimodal data and intelligent systems to automate plant care, optimize resource consumption, and align indoor gardening with sustainable building management practices, paving the way for resilient, green urban spaces.
Why it matches plant phenotyping methods植物の画像・環境センシングと機械学習を統合し、植物形質の抽出および水ストレス推定を行う方法論が研究の中心である。
abstractThis paper presents a novel framework that integrates computer vision, machine learning (ML), and environmental sensing for the automated monitoring of plant health and growth.
As primary producers in Earth's system, plants drive global matter and energy fluxes. Understanding the global distribution of plant functional traits and their biodiversity is, therefore, critical for understanding ecosystem behavior and Earth system dynamics in the face of climate and global change. However, we lack observations for various plant functional traits, such as plant height, leaf size, and nitrogen content, at a global scale.These data gaps could be addressed through citizen science projects, where thousands of individuals have already recorded millions of plant photographs for species identification purposes. While these photographs do not include direct information about plant traits, trait data for thousands of plant species can be accessed from scientific databases. By linking these two data sources—crowd-sourced plant photographs and trait information from scientific databases—through plant species, we can supervise computer vision models to infer plant traits from plant images. The principle of "form follows function" suggests that a plant's appearance can provide valuable insights into its functional properties.To assess the potential of citizen science data for plant trait estimation, we propose testing the feasibility of using weak and noisy labels for effective trait prediction. Considering that different plant traits are not independent of each other, we leverage multi-trait learning. Additionally, our approach incorporates plant images along with ancillary environmental data, such as soil conditions and Earth observation satellite data, to provide crucial context on factors like climate or land surface properties.To fairly evaluate model performance, we curate a clean dataset spanning diverse geographic regions, as well as taxonomic and phylogenetic groups. We conduct a comprehensive study on the resilience of trained models across these distribution shifts. Furthermore, we assess which traits can be effectively learned from noisy labels and explore the extent of trait transferability under different conditions.Our findings indicate that models trained on noisy data can, to a notable extent, predict a series of plant traits, including plant height, leaf area, and specific leaf area. This approach provides an efficient, scalable, and non-destructive method for estimating important plant functional traits. It could lay the groundwork for large-scale biodiversity monitoring and ecosystem assessment, with the potential to revolutionize how we track the functional properties of ecosystems at a global scale.
Why it matches plant phenotyping methods市民科学画像から植物形質を推定するコンピュータビジョン手法を開発・評価し、データセット構築と分布シフト下での性能検証も行っており、植物フェノタイピング手法が中心である。
abstractwe can supervise computer vision models to infer plant traits from plant images
Field / plotMultimodalClassificationYield / biomass estimationBiomass / plant weightGrowth / development / phenologyWater status / transpirationYield / yield components
Climate-smart agriculture aims to implement a suite of conservation management practices, such as cover crops, reduced tillage, smart irrigation and crop rotations, to maximize agroecosystem productivity and reduce greenhouse gas emissions. Timely and high-resolution agriculture data are crucial for measuring, reporting and verifying the implementation and benefits of climate-smart agriculture practices. However, agricultural data collection through field sampling, laboratory analysis, and/or grower surveys is time-consuming and costly. To address these challenges, we developed an artificial intelligence-empowered cross-scale sensing framework to integrate multi-source ground truth data with multi-modal satellite Earth observations to quantify high spatial and temporal information of essential agroecosystem variables in the EU. Specifically, these essential variables include crop types, harvest time, tillage practices, cover crop adoption and biomass, crop yield, soil moisture, ecosystem gross primary productivity and evapotranspiration. We developed computer vision and machine learning algorithms to obtain ground truth data from in-situ measurements, citizen sciences, census surveys, and ground or aerial vehicle system data. Through process-guided machine learning (PGML), we integrated the domain knowledge of soil-vegetation radiative transfer models and ground truth data to accurately quantify these essential variables from Sentinel-1, 2, 3 and SMAP satellite data. This study highlights the potential of integrating cross-scale sensing and PGML to quantify essential ecosystem variables to support climate-smart agriculture.
Why it matches plant phenotyping methods衛星・地上観測と機械学習を統合し、作物バイオマスや収量などの植物関連形質を推定する横断的センシング基盤を開発しており、形質取得・推定手法が中心である。
abstractwe developed an artificial intelligence-empowered cross-scale sensing framework to integrate multi-source ground truth data with multi-modal satellite Earth observations to quantify high spatial and temporal information of essential agroecosystem variables in the EU.
Biomass carbon sequestration and sink capacities of tropical rainforests are vital for addressing climate change. However, canopy height must be accurately estimated to determine carbon sink potential and implement effective forest management. Four advanced machine-learning algorithms—random forest (RF), gradient boosting decision tree, convolutional neural network, and backpropagation neural network—were compared in terms of forest canopy height in the Hainan Tropical Rainforest National Park. A total of 140 field survey plots and 315 unmanned aerial vehicle photogrammetry plots, along with multi-modal remote sensing datasets (including GEDI and ICESat-2 satellite-carried LiDAR data, Landsat images, and environmental information) were used to validate forest canopy height from 2003 to 2023. The results showed that RH80 was the optimal choice for the prediction model regarding percentile selection, and the RF algorithm exhibited the optimal performance in terms of accuracy and stability, with R2 values of 0.71 and 0.60 for the training and testing sets, respectively, and a relative root mean square error of 21.36%. The RH80 percentile model using the RF algorithm was employed to estimate the forest canopy height distribution in the Hainan Tropical Rainforest National Park from 2003 to 2023, and the canopy heights of five forest types (tropical lowland rainforests, tropical montane cloud forests, tropical seasonal rainforests, tropical montane rainforests, and tropical coniferous forests) were calculated. The study found that from 2003 to 2023, the canopy height in the Hainan Tropical Rainforest National Park showed an overall increasing trend, ranging from 2.95 to 22.02 m. The tropical montane cloud forest had the highest average canopy height, while the tropical seasonal forest exhibited the fastest growth. The findings provide valuable insights for a deeper understanding of the growth dynamics of tropical rainforests.
Why it matches plant phenotyping methods森林キャノピー高という植物群落の形態形質を対象に、複数の機械学習・マルチモーダルリモートセンシング手法を比較し、現地調査およびUAVデータで検証しているため、形質推定法が中心である。
abstractFour advanced machine-learning algorithms—random forest (RF), gradient boosting decision tree, convolutional neural network, and backpropagation neural network—were compared in terms of forest canopy height
A novel eggplant disease detection method based on multimodal data fusion and attention mechanisms is proposed in this study, aimed at improving both the accuracy and robustness of disease detection. The method integrates image and sensor data, optimizing the fusion of multimodal features through an embedded attention mechanism, which enhances the model's ability to focus on disease-related features. Experimental results demonstrate that the proposed method excels across various evaluation metrics, achieving a precision of 0.94, recall of 0.90, accuracy of 0.92, and mAP@75 of 0.91, indicating excellent classification accuracy and object localization capability. Further experiments, through ablation studies, evaluated the impact of different attention mechanisms and loss functions on model performance, all of which showed superior performance for the proposed approach. The multimodal data fusion combined with the embedded attention mechanism effectively enhances the accuracy and robustness of the eggplant disease detection model, making it highly suitable for complex disease identification tasks and demonstrating significant potential for widespread application.
Why it matches plant phenotyping methodsナスの病害状態を画像・センサーデータから推定する手法を開発し、アブレーション実験と性能評価で検証しているため、植物フェノタイピング手法が中心です。
abstractA novel eggplant disease detection method based on multimodal data fusion and attention mechanisms is proposed in this study
Unmanned aerial vehicle (UAV) platforms are increasingly used to obtain plant phenotypes in crop breeding for their efficiency and versatility. A lightweight UAV was used to collect high-precision RGB images, multispectral and point cloud data of soybeans (Glycine max (L.) Merr.) across fields at various growth stages, utilizing an innovative cross-circling oblique (CCO) route. A multi-modal data fusion deep learning model was proposed based on the self-supervised contrastive learning strategy with fine-tuning for yield estimation and lodging discrimination in soybean germplasm resources. During the soybean growth stages of flowering (R1) to maturity (R8), the contrastive learning effectively captured the decoupling characteristics of different soybean varieties in the feature space. Higher accuracy in yield estimation was obtained combined contrastive learning with the traditional features. Correlations were significantly reduced between features among varieties (Pearson’s mean 0.27–0.62) and feature separations were achieved after dimension reduction (R8: CH = 12.4, DB = 51.8). RMSE of yield estimation was 591.39 kg ha⁻¹ at high density and 532.75 kg ha⁻¹ at low density at R8 growth stages. Lodging discrimination achieved the highest accuracy with an F1-score of 0.57 at high density and 0.64 at low density. The results demonstrated that utilizing contrastive learning for extraction of deep soybean features holds significant potential in supporting traditional features for yield estimation and lodging discrimination.
Why it matches plant phenotyping methodsUAVによるマルチモーダル植物表現型取得と、対照学習を用いた収量推定・倒伏判別モデルが研究の中心であり、手法性能も定量評価している。
abstractA multi-modal data fusion deep learning model was proposed based on the self-supervised contrastive learning strategy with fine-tuning for yield estimation and lodging discrimination in soybean germplasm resources.
Stomatal conductance (g s ) quantifies the rate of exchange of carbon dioxide for photosynthesis and water vapor for transpiration between plant leaves and the atmosphere. g s is usually measured by handheld devices like porometers , and readings are manually taken in the field, which is time-consuming and labor-intensive. In this study, we investigated the use of high-throughput phenotyping (HTP) data combined with weather data to estimate g s through machine-learning (ML) modeling. The experiment was conducted in a research field equipped with an HTP platform in 2020 and 2021 involving maize, sorghum, soybean, sunflower , and winter wheat . Weather variables including dew point temperature, wind speed , air temperature, solar radiation, and relative humidity were collected by an onsite weather station . Plot-level canopy temperature, soil temperature , and seven vegetation indices were acquired using a thermal infrared camera, a multispectral camera, and a visible near-infrared spectrometer integrated on the HTP platform. Three supervised ML methods (Partial Least Squares Regression (PLSR), Random Forest Regression (RFR), and Support Vector Regression (SVR)) were employed to train the estimation models for g s , and model performance was evaluated by Coefficient of Determination (R 2 ) and Root Mean Squared Error (RMSE). The result showed that RFR and SVR outperformed PLSR in g s modeling. The RFR model achieved R 2 of 0.63 and RMSE of 0.16 mol m −2 ·s −1 with the combination of phenotyping data and weather data. It outperformed the model using only the weather data (R 2 =0.35 and RMSE=0.21 mol m −2 ·s −1 ), or the model using only the phenotyping data (R 2 =0.46 and RMSE=0.19 mol m −2 ·s −1 ). This result suggested that high-throughput plant phenotyping data effectively complement weather data in estimating g s rapidly and non-destructively through ML. With the wide adoption of HTP technologies in aerial and ground-based platforms, this research provides a practical framework to estimate g s at large scale for crop breeding and irrigation management .
Why it matches plant phenotyping methodsHTPセンサーデータと機械学習を用いて、植物の生理形質である気孔コンダクタンスを大規模・非破壊推定する方法が研究の中心であり、モデル性能も評価している。
abstractIn this study, we investigated the use of high-throughput phenotyping (HTP) data combined with weather data to estimate g s through machine-learning (ML) modeling.
Corn is one of the important food crops and industrial raw materials. However, maize diseases have seriously affected its yield and quality. In order to effectively identify maize diseases, digital image processing technology has been widely used in the agricultural field. The classification of diseases based on digital images enables early detection of corn diseases, reducing farmers’ losses. Although existing corn disease identification methods have made significant progress using deep learning technology for digital image processing, most of these technologies rely on single-modal data for identification and lack the connection between images and texts. To solve this problem, this paper proposes a cross-modal feature alignment fusion model called WCG-VMamba. Firstly, we propose a wavelet visual Mamba (WAVM) network, which integrates the advantages of different visual coding strategies and can reduce the influence of intrinsic noise and other factors on the validity of image features during extraction. Then, we introduce the Cross Modal Alignment Transformer (CMAT), which interacts image features with text features to capture their semantic correlation and determine the weight distribution of image and text features in the fusion process. We then use Transformer coding blocks to fuse features. Finally, the Gaussian Random Walk Duck Swarm Algorithm (GRW-DSA) is proposed to reduce errors in the duck swarm exploration process through Gaussian Random Walk, aiming to find the optimal learning rate. Experiments on self-built datasets and two common datasets show that WCG-VMamba can be effectively used in the task of corn disease classification. Compared with other excellent models such as MobileViT, MobilenetV3, SwinT, and DINOV2, better results have been achieved, our model achieves a recognition accuracy as high as 96.97%, proving its important practical application in promoting agricultural cross-modal models and corn disease control.
Why it matches plant phenotyping methodsトウモロコシ葉の画像から病害を分類する新規マルチモーダルモデルを開発し、複数データセットで比較評価している。植物の病害状態を直接推定する手法が研究の中心である。
abstractthis paper proposes a cross-modal feature alignment fusion model called WCG-VMamba
It is difficult and costly to obtain large-scale, labeled crop disease data in the field of agriculture. How to use small samples of unlabeled data for feature learning has become an urgent problem that needs to be solved. The emergence of self-supervised contrastive learning methods and self-supervised mask learning methods can solve the problem of missing labels on the training data. However, each of these paradigms comes with its own advantages and drawbacks. At the same time, the features learned by dataset in a single modality are limited, ignoring the correlation with other modal information. Hence, this paper introduced an effective framework for multimodal self-supervised learning, denoted as MMSSL, to address the task of identifying cucumber diseases with small sample sizes. Integrating image self-supervised mask learning, image self-supervised contrastive learning, and multimodal image-text contrastive learning, the model can not only learn disease feature information from different modalities, but also capture global and local disease feature information. Simultaneously, the mask learning branch was enhanced by introducing a prompt learning module based on a cross-attention network. This module aided in approximately locating the masked regions in the image data in advance, facilitating the decoder in making accurate decoding predictions. Experimental results demonstrate that the proposed method achieves a 95% accuracy in cucumber disease identification in the absence of labels. The approach effectively uncovers high-level semantic features within multimodal small-sample cucumber disease data. GradCAM is also employed for visual analysis to further understand the decision-making process of the model in disease identification. In conclusion, the proposed method in this paper is advantageous for enhancing the classification accuracy of small-sample cucumber data in a multimodal, unlabeled context, demonstrating good generalization performance.
Why it matches plant phenotyping methodsキュウリ病害の画像から植物の病徴状態を推定するマルチモーダル自己教師あり学習法を開発・評価しており、病害識別という植物フェノタイプ推定が中心的な技術貢献である。
abstractthis paper introduced an effective framework for multimodal self-supervised learning, denoted as MMSSL, to address the task of identifying cucumber diseases with small sample sizes.
Temperate mixed forest ecosystems consist of diverse plant functional types (PFTs) that exhibit variations in phenology and physiological responses to climate change. Consequently, the traditional big-leaf assumptions in carbon modeling have been criticized for oversimplifying these ecosystems, for they overlook the variability in PFT composition and their sensitivity to climate within these ecosystems. However, incorporating PFT composition into carbon and climate sensitivity simulations in heterogeneous mixed forest ecosystems presents two major challenges: (1) accurate fine-scale PFT composition mapping across large forest landscapes remains lacking, which further leads to (2) incomplete assessments of these fine-scale PFT contributions in interpreting ecosystem-scale carbon dynamics and climate sensitivity response. The recent increase in high-resolution satellite and ground observation data offers an unprecedented opportunity to resolve these challenges.To address the first challenge, we developed a novel approach integrating Fisher-transformation-based unmixing analysis with time-series spectral and radar data. We examined this approach in three representative temperate mixed landscapes in the northeastern United States, using time-series Sentinel-1 and -2 data for calibration and local airborne-derived PFT fraction maps for validation. Our results demonstrate that (1) the synergy of spectral and radar time-series features significantly improves accuracy compared to spectral time-series models; (2) optimized features based on the Fisher-transformation approach minimize within-PFT variability and maximize between-PFT variability, enhancing model generalizability across landscapes. Integration of this approach with Google Earth Engine enables accurate ecoregion-wise PFT fractional mapping. To address the second challenge, we integrate different levels of PFT-related characteristics (e.g., PFT fraction map, PFT-specific physiology approximated by satellite vegetation index) with a machine learning-based carbon modeling scheme, examining how these PFT-related characteristics and climate variables separately and jointly determined the net ecosystem carbon exchange (NEE) in real mixed forest ecosystems. Specifically, we used the CHEESEHEAD19 dataset, which includes the world's most densely distributed eddy-covariance (EC) flux towers (13+ towers) within a 10 km × 10 km domain in the mixed forest ecoregion of Wisconsin, US, providing half-hourly flux records. Daily, 3-meter resolution, gap-free maps of vegetation index (NIRv) were calculated using the PlanetScope surface reflectance product. Our results demonstrated that PFT-related characteristics play a significant role (˜50%) in interpreting half-hourly NEE dynamics, with PFT-specific NIRv playing a dominant role (~30%), followed by PFT-fraction (~20%). Furthermore, by partitioning PFT-related effects, our results reveal distinct NEE sensitivity responses to specific environmental variability within and between PFTs in the 10 km × 10 km forest landscapes. Collectively this work advances the mapping of PFT composition and highlights the importance of integrating these fine-scale forest compositions into carbon modelling and climate sensitivity assessments, particularly in heterogeneous temperate mixed ecosystems.
Why it matches plant phenotyping methods衛星スペクトル・レーダー時系列から植物機能型(PFT)構成比を推定する手法を開発し、航空機由来マップで検証しており、植物キャノピー状態の取得・推定が中心的な技術貢献である。
abstractwe developed a novel approach integrating Fisher-transformation-based unmixing analysis with time-series spectral and radar data
Maize leaf area offers valuable insights into physiological processes, playing a critical role in breeding and guiding agricultural practices. The Azure Kinect DK possesses the real-time capability to capture and analyze the spatial structural features of crops. However, its further application in maize leaf area measurement is constrained by RGB–depth misalignment and limited sensitivity to detailed organ-level features. This study proposed a novel approach to address and optimize the limitations of the Azure Kinect DK through the multimodal coupling of RGB-D data for enhanced organ-level crop phenotyping. To correct RGB–depth misalignment, a unified recalibration method was developed to ensure accurate alignment between RGB and depth data. Furthermore, a semantic information-guided depth inpainting method was proposed, designed to repair void and flying pixels commonly observed in Azure Kinect DK outputs. The semantic information was extracted using a joint YOLOv11-SAM2 model, which utilizes supervised object recognition prompts and advanced visual large models to achieve precise RGB image semantic parsing with minimal manual input. An efficient pixel filter-based depth inpainting algorithm was then designed to inpaint void and flying pixels and restore consistent, high-confidence depth values within semantic regions. A validation of this approach through leaf area measurements in practical maize field applications—challenged by a limited workspace, constrained viewpoints, and environmental variability—demonstrated near-laboratory precision, achieving an MAPE of 6.549%, RMSE of 4.114 cm2, MAE of 2.980 cm2, and R2 of 0.976 across 60 maize leaf samples. By focusing processing efforts on the image level rather than directly on 3D point clouds, this approach markedly enhanced both efficiency and accuracy with the sufficient utilization of the Azure Kinect DK, making it a promising solution for high-throughput 3D crop phenotyping.
Why it matches plant phenotyping methodsAzure Kinect RGB-D再較正、深度補間、意味解析を開発し、トウモロコシ葉面積測定で検証した、植物表現型取得法が中心の研究。
abstractThis study proposed a novel approach to address and optimize the limitations of the Azure Kinect DK through the multimodal coupling of RGB-D data for enhanced organ-level crop phenotyping.
Accurate background segmentation in 3D plant phenotyping is crucial for reliable trait assessment but remains challenging. Current methods are either excessively complex, developed for a different domain, or lead to data loss (coordinate-based). This paper addresses these issues by introducing an AI-driven approach using a Multi-Layer Perceptron (MLP) model, leveraging RGB, spatial (XYZ), and near-infrared (NIR) data to enhance precision. The method was evaluated on high-throughput phenotyping data, achieving a classification accuracy of 0.9993, significantly reducing false positives and false negatives compared to coordinate-based segmentation. The proposed approach improved segmentation, particularly in early growth stages and for prostrate species, where traditional methods often fail. The model’s impact on leaf area estimation was validated against destructive measurements, demonstrating substantial accuracy improvements, especially for species with small and prostrate canopies. Additionally, the model exhibited strong generalization capabilities when applied to an external 3D dataset, confirming its reusability beyond plant phenotyping tasks. Integrating this simple method into phenotyping pipelines will enhance efficiency and accuracy in high-throughput trait estimation, supporting advancements in plant science and precision agriculture.
Why it matches plant phenotyping methods3D植物フェノタイピング用の背景セグメンテーション手法を開発し、葉面積推定を破壊的測定および外部データセットで検証しており、表現型取得・抽出手法が中心である。
abstractThis paper addresses these issues by introducing an AI-driven approach using a Multi-Layer Perceptron (MLP) model, leveraging RGB, spatial (XYZ), and near-infrared (NIR) data to enhance precision.
Nitrogen use efficiency (NUE) is a key indicator for selecting nitrogen-efficient crop cultivars and optimizing fertilization strategies. However, NUE is typically assessed using destructive and laborious sampling methods, hindering the advancement of sustainable agriculture. The objective is to test whether the fusion of phenotyping data simultaneously acquired by multi-source sensors that reflect more functional and structural traits can improve the estimation accuracy of the highly integrated trait NUE in maize. Multispectral (MS) and light detection and ranging (LiDAR) data were simultaneously acquired during critical growth stages across two years of maize cultivar and nitrogen fertilizer field experiments using a multi-sensor UAV platform. Three machine learning algorithms, Partial Least Squares Regression (PLSR), Random Forest Regression (RFR) and Support Vector Machine Regression (SVR) were selected to construct NUE estimation models based on three data sources:MS, LiDAR, and MS+LiDAR. The results demonstrated distinct differences in nitrogen utilization efficiency (NUₜE) and nitrogen agronomy efficiency (NAE) among maize cultivars at critical growth stages. These differences were efficiently and accurately identified using multi-source data combined with the machine learning algorithms. The RFR method obtained the highest model validation accuracy with an average Rₜₑₛₜ² = 0.68 and RMSEₜₑₛₜ = 6.66kgkg⁻¹. The average accuracy of multi-source data fusion was improved by 20.21% compared to a single data source, and the RFR+MS+LiDAR method for NUₜE estimation obtained the highest model accuracy in the two-year validation dataset with Rₜₑₛₜ² = 0. 86 and RMSEₜₑₛₜ = 8.5kgkg⁻¹. The method proposed in this study mitigates the impact of canopy spectral saturation during the late growth stages of maize, enhancing the accuracy of NUE estimation by improving the convergence between predicted and observed values. This multi-source data fusion approach, based on a UAV platform, enables effective monitoring of NUE at critical growth stages. Consequently, it advances rapid, non-destructive NUE assessment in maize, supporting efficient breeding and precision nitrogen management strategies.
Why it matches plant phenotyping methodsUAVマルチスペクトル画像とLiDARを融合し、機械学習でトウモロコシの窒素利用効率という植物形質を非破壊推定する手法を開発・検証しており、フェノタイピング手法が中心である。
abstractThe objective is to test whether the fusion of phenotyping data simultaneously acquired by multi-source sensors that reflect more functional and structural traits can improve the estimation accuracy of the highly integrated trait NUE in maize.
Powdery mildew disease threatens wheat production worldwide, and early detection is of great significance for disease control and maximizing yield and quality. To improve early remote sensing detection of wheat powdery mildew, solar-induced chlorophyll fluorescence (SIF) parameters were extracted using three-band Fraunhofer line discrimination (3FLD) and reflectance index approaches, and vegetation index (VI) was calculated by hyperspectral reflectance. All features and feature subsets of different data sources were used as inputs to multiple linear regression (MLR), random forest (RF), and support vector machine (SVM) algorithms to construct a wheat powdery mildew monitoring model. SVM includes linear kernel function (LK), polynomial kernel function (PK), and Gaussian radial basis function (RBF). Under wheat powdery mildew stress, wheat canopy reflectance showed a blue shift, and fluorescence weakened. The correlation between SIF−A intensity and disease index (DI) in the O²−A band extracted using the 3FLD method was the highest at −0.781, showing that the SIF parameter was useful for monitoring powdery mildew. Whether based on all features or feature subsets, the RBF model achieved the highest model accuracy, followed by the RF and the MLR. In the feature subset, the accuracy ranges of RBF, LK, and PK models are 0.740−0.871, 0.724−0.850, and 0.716−0.841 respectively. The SIF+VI in the RBF model is more useful for early and stable disease monitoring of wheat powdery mildew. This innovative technical solution is expected to support the early diagnosis of wheat powdery mildew, significantly improving disease prevention and control efficiency and effectiveness.
Why it matches plant phenotyping methods小麦の病害状態をSIF・ハイパースペクトル反射から推定する検出手法とモデルを開発・評価しており、植物表現型取得が中心である。
abstractTo improve early remote sensing detection of wheat powdery mildew, solar-induced chlorophyll fluorescence (SIF) parameters were extracted using three-band Fraunhofer line discrimination (3FLD) and reflectance index approaches, and vegetation index (VI) was calculated by hyperspectral reflectance.
The study of sodium content in plants is crucial for the improvement of saline-alkali soil. Existing metal element detection methods pose challenges because they are complicated and time-consuming. In this study, we propose a quantitative detection model, FusionNet, that integrates Laser-Induced Breakdown Spectroscopy (LIBS) and Near-Infrared Hyperspectral Imaging (NIR-HSI) to realize the detection of Na element content in sorghum roots. To address the small-sample dataset, A Generative Adversarial Network (GAN) was employed to increase the diversity of the samples. The results indicated that data augmentation effectively enhanced the diversity of the original dataset and improved model performance. The modeling results from the FusionNet network achieved R²cv of 0.9915 and RMSECV of 0.7418, while R²p and RMSEP were 0.9808 and 0.6693. Compared to training with LIBS data alone, FusionNet achieved improvements of 4.94 % in R²cv and 5.61 % in R²p. This study provides a new method for detecting metal elements in plants.
Why it matches plant phenotyping methods植物根のNa含量という形質を、LIBSとNIR-HSIのデータ融合およびFusionNetで定量推定する手法を開発・評価しており、表現型取得が研究の中心である。
abstractwe propose a quantitative detection model, FusionNet, that integrates Laser-Induced Breakdown Spectroscopy (LIBS) and Near-Infrared Hyperspectral Imaging (NIR-HSI) to realize the detection of Na element content in sorghum roots.
The aboveground biomass (AGB) of crops is an essential metric for monitoring crop growth, making timely and accurate AGB forecasting critical for effective agricultural management. The introduction of Unmanned Aerial Vehicles (UAVs) and advanced sensor technologies has revolutionized traditional AGB prediction techniques. Currently, machine learning (ML) combined with UAV data are commonly utilized, along with the Vegetation Index Weighted Canopy Volume Model (CVMVI) for AGB prediction. Nevertheless, there is limited investigation into how these methods perform across different agricultural conditions. This study aims to fill this gap by creating specific methodologies for estimating corn AGB under diverse fertilization and irrigation treatments. We utilized LiDAR, multispectral (MS), thermal infrared (TIR), along with measured AGB and Leaf Area Index (LAI) data from various growth stages to develop a stacking ensemble learning model. This model effectively integrates data from multiple sources, resulting in a strong prediction performance with R² of 0.86, Mean Absolute Error (MAE) of 1.54 t/ha, and Root Mean Square Error (RMSE) of 2.06 t/ha. Meanwhile, the analysis of the accuracy of CVMVI revealed its efficacy during the early-stage when corn is short, with its predictive capability diminishing as AGB increases. Consequently, we recommend the CVMVI for early-stage AGB prediction, which can streamline data collection and computational efforts. In contrast, the ML approach, which benefits from data fusion, is more appropriate for predicting AGB during the mid to late growth stages. This study enhances AGB prediction accuracy and speed, providing critical understanding of regional AGB dynamics and supporting better agricultural decision-making.
Why it matches plant phenotyping methodsUAVのLiDAR・マルチスペクトル・熱赤外データを用いてトウモロコシの地上部バイオマスを推定する機械学習モデルと体積モデルを開発・比較し、精度を評価しているため、植物形質取得手法が研究の中心である。
abstractThis study aims to fill this gap by creating specific methodologies for estimating corn AGB under diverse fertilization and irrigation treatments.
In recent years, the rapid advancement of computer vision and machine learning techniques has revolutionized the field of agriculture by enabling automated detection and classification of plant diseases. This work presents a comprehensive approach for the efficient segmentation and multi-class classification of heterogeneous plant disease and non-disease images. The proposed methodology combines multi-level thresholding-based filtering, feature ranking, leaf feature segmentation, and ensemble classification to achieve high accuracy in disease recognition. In the initial stage, multi-level thresholding-based filtering is applied to enhance image quality and reduce noise, facilitating more precise disease pattern extraction. Subsequently, feature ranking methods are employed to select the most discriminative features from the pre-processed images. These features play a pivotal role in distinguishing between various plant diseases and healthy leaves. Experimental results on a multi-modal plant disease and non- disease images demonstrate the superiority of our proposed method in terms of accuracy and efficiency.
Why it matches plant phenotyping methods植物病害画像から葉の病害状態を抽出・分類する画像解析手法が研究の中心であり、単なる生物学的実験の測定ではないため。
abstractThis work presents a comprehensive approach for the efficient segmentation and multi-class classification of heterogeneous plant disease and non-disease images.
Dissecting the drought resistance (DR) mechanism and designing drought-resistant rice varieties are promising strategies to address the challenge of climate change. Here, we selected a typical drought-avoidant (DA) variety IRAT109 and drought-tolerant (DT) variety Hanhui15 as the parents to develop a stable recombinant inbred line (RIL) population (F 8 , 1,262 lines). The de novo assembled genomes of both parents were released. Through re-sequencing of the RIL population, a set of 1,189,216 reliable SNPs were obtained and used for constructing a dense genetic map. Using both aboveground and underground phenomic platforms and multimodal cameras, we captured 139,040 image-based traits (i-traits) of whole plant’s phenotypes in response to drought stress throughout entire rice growth period and identified 32,586 drought-responsive quantitative trait loci (QTLs) including 2,097 unique QTLs. The QTLs related to panicle i-traits occurred on the middle of chromosome 8 over 600 times, while the QTLs related to leaf i-traits on the 5’ end of chromosome 3 over 800 times, indicating potential effect of these QTLs on plant phenotypes. We chose three candidate genes ( OsMADS50, OsGhd8, OsSAUR11 ) related to leaf, panicle, and root traits respectively and verified their functions in resisting drought. Gene OsMADS50 was found to negatively regulate DR by modulating leaf dehydration, grain size, and root downward growth. Furthermore, a total of 18 and 21 composite QTLs significantly related to grain weight and plant biomass were screened from 597 lines in RIL population under drought conditions in field experiments, and composite QTL region was highly overlapped (76.9%) with known DR gene region. Based on three candidate DR genes, we proposed the haplotype design suitable for different environments and breeding objectives. This study provides a valuable reference for multi-modal and time-series phenomic analyses, deciphers the genetic mechanism of DA and DT rice varieties, and offers a molecular navigation map for breeding DR variety.
Why it matches plant phenotyping methods地下・地上フェノミックプラットフォームとマルチモーダルカメラで全生育期間の画像形質を大量取得しており、フェノタイピング手法の適用と技術的ワークフローが研究の中核です。
abstractUsing both aboveground and underground phenomic platforms and multimodal cameras, we captured 139,040 image-based traits (i-traits) of whole plant’s phenotypes in response to drought stress throughout entire rice growth period
Reproduction assets foundThe paper's phenome data (aboveground and belowground rice images/i-traits) and the authors' data-handling code and deep-learning model are explicitly deposited at public URLs listed in the Data Availability Statement. Genome data (riceome.hzau.edu.cn) is molecular omics and excluded.Code · publicAll the phenome data and core data-handling code have been deposited online.Open asset ↗lines:140-175Plant phenotyping relevance match · UnverifiedEurope PMC · checked 15 Sept 2026
Published1 Dec 2024Computers and Electronics in Agriculture.
Faba bean is a global food legume crop, and it is essential to accurately and timely determine its plant height, above-ground biomass (fresh and dry weight) and yield for enhancing cultivation practices and planning the next planting season. Traditional ground sampling is a time-consuming and labor-intensive approach. However, the utilization of an unmanned aerial vehicle (UAV) as a high-throughput technique offers a promising alternative strategy for estimating crop phenotypic traits. In this study, a two-year experiment was conducted from 2020 to 2022, where UAV-based multimodal data were collected using red-green-blue, multispectral and thermal infrared sensors. The variables derived from these three sensors and their combinations were used to estimate the fresh weight, dry weight and yield of faba bean based on extreme gradient boosting (XGBoost), random forest, multiple linear regression and k-nearest neighbor algorithms. The following findings were obtained: (1) The use of the maximum percentile crop surface model resulted in the highest estimation accuracy for faba bean plant height. (2) Fusion data from multiple sensors increased the estimation accuracy of faba bean fresh weight, dry weight and yield, the coefficient of determination (R²) improved by 14.22%, 1.45%, and 18.76%, respectively, compared with the best estimation accuracy of a single sensor. (3) The XGBoost algorithm outperformed the other algorithms in estimating fresh weight, dry weight and yield of faba bean. These results demonstrate that multiple sensors and appropriate algorithms can be used to effectively estimate faba bean phenotypic traits and provide valuable insights for agricultural remote sensing research.
Why it matches plant phenotyping methodsUAVマルチモーダルセンサーと機械学習を用いて、ソラマメの草高・バイオマス・収量を推定する手法が研究の中心であり、技術的な比較・精度評価も行っている。
abstractThe variables derived from these three sensors and their combinations were used to estimate the fresh weight, dry weight and yield of faba bean based on extreme gradient boosting (XGBoost), random forest, multiple linear regression and k-nearest neighbor algorithms.
With the development of deep learning technology, the control of tomato diseases has emerged as a crucial aspect of intelligent agricultural management. While current research on tomato disease segmentation has made considerable strides, challenges persist due to the susceptibility of tomato leaf diseases to strong light reflections and shadow gradients in sunlight. Additionally, the complex backgrounds found in agricultural fields often lead to model confusion, resulting in inaccurate segmentation. Traditional methods for tomato disease segmentation rely on single-modal image-based models, which struggle when dealing with the nuanced features and limited scope of tomato leaf diseases. To address these issues, our study introduces the LVF framework, a dual-modal approach combining image and text information for pre-segmentation of tomato diseases. We began by creating a new dataset labeled with both images and text, specifically focusing on diseased tomato leaves with guidance from agricultural experts. For image processing, we developed a probabilistic differential fusion network to mitigate interference caused by high-frequency noise, leveraging color and grayscale images. Furthermore, our reinforcement feature network and threshold filtering network enhance useful information while filtering out negative information from the fused images. In text processing, we proposed a multi-scale cross-nesting network to integrate semantic information about diseases across different scales and types. By nesting Bert-processed word vectors with fused image vectors, our model gains a deeper understanding of semantic information, thereby improving its ability to segment crop diseases accurately. Our experiments, conducted on self-constructed tomato datasets as well as public datasets for tomatoes and maize, demonstrated the efficacy and robustness of our approach in leaf disease segmentation. The LVF framework offers a valuable tool to enhance the accuracy of crop disease segmentation, especially in complex agricultural environments.
Why it matches plant phenotyping methodsトマト葉の病害症状を画像からセグメンテーションする手法を開発し、独自データセットの構築と公開データセットでの評価を行っており、植物状態の取得・抽出が研究の中心である。
abstractour study introduces the LVF framework, a dual-modal approach combining image and text information for pre-segmentation of tomato diseases
The increasing frequency of typhoon events, attributed to global climate change, has significantly affected agricultural production, predominantly resulting in substantial negative consequences. Accurate and timely assessment of crop damage is crucial for understanding economic implications, devising effective agricultural strategies, and enhancing resilience amid mounting climate uncertainties. This study investigates the utility of Sentinel-1 Synthetic Aperture Radar (SAR) and Sentinel-2 MultiSpectral Instrument (MSI) data in tracking maize damage severity following typhoon events. By employing continuous field sampling techniques and conducting visual interpretation of high-resolution remote sensing imagery, sample sets representing a spectrum of maize damage severity were systematically established. The application of time series analysis on Sentinel data enables a comprehensive exploration of spectral and polarization responses, providing insights into the correlation with maize damage severity. Segmentation of the maize damage timeline into pre-disaster, disaster, and recovery periods, coupled with optimization of relevant feature parameters, was undertaken to bolster monitoring precision. Leveraging the Google Earth Engine (GEE) cloud platform, a Random Forest algorithm was used to develop a model for monitoring maize damage severity across different post-typhoon periods, yielding maps delineating the distribution and magnitude of maize damage in Northeast China. Results indicate that integrating spectral indices from the pre-disaster phase with backscatter variations of polarization bands during various post-typhoon periods enhances maize damage assessment. Maize damage severity is notably elevated during the disaster period, achieving an overall accuracy of 87.22 %. While mitigated during the recovery phase, localized exacerbation occurs in severely affected regions, yielding an overall accuracy of 88.54 %. Analysis incorporating terrain and meteorological data reveals that post-typhoon maize disasters predominantly occur in low-lying and flat areas, with meteorological factors, particularly maximum wind speed and daily cumulative precipitation, exerting significant influence on damage severity. This study underscores the critical role of SAR and optical data fusion in elucidating typhoon-induced crop damage dynamics, thereby providing essential insights for proactive mitigation strategies against future agricultural losses.
Why it matches plant phenotyping methodsSentinel-1/2データとランダムフォレストを用いて、トウモロコシの台風被害 severity という植物状態を推定・検証する手法が研究の中心である。
abstractThis study investigates the utility of Sentinel-1 Synthetic Aperture Radar (SAR) and Sentinel-2 MultiSpectral Instrument (MSI) data in tracking maize damage severity following typhoon events.
Nitrogen (N) nutrition index (NNI) is a reliable indicator of plant N status for field crops, but its determination is both labor- and cost-intensive. The utilization of remote sensing approaches for monitoring N, mainly in relevant crops such as of corn (Zea mays L.), will be critical for enhancing effective use of this nutrient. Therefore, the aim of this study was to assess NNI predicted from optical and C-band Synthetic Aperture Radar (C-SAR) satellite data and available soil N (Nₐᵥ) at different vegetative growth stages for corn crop. Eleven field studies were conducted in the Pampas region (Argentina), applying five fertilizer N rates (0, 60, 120, 180, and 240 kg N ha⁻¹), all at sowing time. Plant samples were collected at sixth-leaf (V₆), tenth-leaf (V₁₀), fourteen-leaf (V₁₄), and flowering (R₁). Using linear regression models, NNI was best predicted using only optical satellite data from V₆ to V₁₄, and integrating optical with C-SAR plus Nₐᵥ at R₁. The best monitoring model integrated vegetation spectral indices, C-SAR and Nₐᵥ data at V₁₀ with an adjusted R² of 0.75 achieved during calibration in the northern Pampa. During validation, it predicted NNI with an RMSE of 0.14 and a MAPE of 12% in the southeastern Pampa. The red-edge spectrum and Local Incidence Angle of C-SAR were necessary to monitor the corn N status via prediction of NNI. Thus, this study provided empirical models to remotely sensed corn N status within fields during vegetative period, serving as a foundational data for guiding future N management.
Why it matches plant phenotyping methods衛星光学・C-SARデータからトウモロコシの窒素栄養指数を推定するモデルを開発・検証しており、植物形質推定が研究の中心である。
abstractThe utilization of remote sensing approaches for monitoring N, mainly in relevant crops such as of corn (Zea mays L.), will be critical for enhancing effective use of this nutrient.
Traditional methods for detecting seed germination rates often involve lengthy experiments that result in damaged seeds. This study selected the Zheng Dan-958 maize variety to predict germination rates using multi-source information fusion and a random forest (RF) algorithm. Images of the seeds and internal cracks were captured with a digital camera. In contrast, the dielectric constant of the seeds was measured using a flat capacitor and converted into voltage readings. Features such as color, shape, texture, crack count, and normalized voltage were used to form feature vectors. Various prediction algorithms, including random forest (RF), radial basis function (RBF), neural networks (NNs), support vector machine (SVM), and extreme learning machine (ELM), were developed and tested against standard germination experiments. The RF model stood out, with a training time of 5.18 s and the highest accuracy of 92.88%, along with a mean absolute error (MAE) of 0.913 and a root mean square error (RMSE) of 1.163. The study concluded that the RF model, combined with multi-source information fusion, offers a feasible and nondestructive method for quickly and accurately predicting maize seed germination rates.
Why it matches plant phenotyping methods種子画像と誘電計測を融合し、発芽率という植物状態を非破壊推定する手法を開発・比較検証しており、表現型取得・推定法が研究の中心である。
abstractThis study selected the Zheng Dan-958 maize variety to predict germination rates using multi-source information fusion and a random forest (RF) algorithm.
Plant phenotyping relevance match · UnverifiedOpenAlex · Europe PMC · checked 15 Sept 2026
Abstract Background The early and specific detection of abiotic and biotic stresses, particularly their combinations, is a major challenge for maintaining and increasing plant productivity in sustainable agriculture under changing environmental conditions. Optical imaging techniques enable cost-efficient and non-destructive quantification of plant stress states. Monomodal detection of certain stressors is usually based on non-specific/indirect features and therefore is commonly limited in their cross-specificity to other stressors. The fusion of multi-domain sensor systems can provide more potentially discriminative features for machine learning models and potentially provide synergistic information to increase cross-specificity in plant disease detection when image data are fused at the pixel level. Results In this study, we demonstrate successful multi-modal image registration of RGB, hyperspectral (HSI) and chlorophyll fluorescence (ChlF) kinetics data at the pixel level for high-throughput phenotyping ofA. thalianagrown in Multi-well plates and an assay with detached leaf discs ofRosa × hybridainoculated with the black spot disease-inducing fungusDiplocarpon rosae. Here, we showcase the effects of (i) selection of reference image selection, (ii) different registrations methods and (iii) frame selection on the performance of image registration via affine transform. In addition, we developed a combined approach for registration methods through NCC-based selection for each file, resulting in a robust and accurate approach that sacrifices computational time. Since image data encompass multiple objects, the initial coarse image registration using a global transformation matrix exhibited heterogeneity across different image regions. By employing an additional fine registration on the object-separated image data, we achieved a high overlap ratio. Specifically, for theA. thalianatest set, the overlap ratios (ORConvex) were 98.0 ± 2.3% for RGB-to-ChlF and 96.6 ± 4.2% for HSI-to-ChlF. For theRosa × hybridatest set, the values were 98.9 ± 0.5% for RGB-to-ChlF and 98.3 ± 1.3% for HSI-to-ChlF. Conclusion The presented multi-modal imaging pipeline enables high-throughput, high-dimensional phenotyping of different plant species with respect to various biotic or abiotic stressors. This paves the way for in-depth studies investigating the correlative relationships of the multi-domain data or the performance enhancement of machine learning models via multi modal image fusion.
Why it matches plant phenotyping methodsRGB・HSI・クロロフィル蛍光画像の画素レベル登録手法と統合パイプラインを開発・評価し、高スループット植物表現型解析に直接利用しているため。
abstractwe demonstrate successful multi-modal image registration of RGB, hyperspectral (HSI) and chlorophyll fluorescence (ChlF) kinetics data at the pixel level for high-throughput phenotyping
Effective lettuce cultivation requires precise monitoring of growth characteristics, quality assessment, and optimal harvest timing. In a recent study, a deep learning model based on multimodal data fusion was developed to estimate lettuce phenotypic traits accurately. A dual-modal network combining RGB and depth images was designed using an open lettuce dataset. The network incorporated both a feature correction module and a feature fusion module, significantly enhancing the performance in object detection, segmentation, and trait estimation. The model demonstrated high accuracy in estimating key traits, including fresh weight (fw), dry weight (dw), plant height (h), canopy diameter (d), and leaf area (la), achieving an R2 of 0.9732 for fresh weight. Robustness and accuracy were further validated through 5-fold cross-validation, offering a promising approach for future crop phenotyping.
Why it matches plant phenotyping methodsRGB・深度画像を融合した深層学習によるレタス形質推定手法を開発し、交差検証で性能評価しており、フェノタイピング手法が中心である。
abstracta deep learning model based on multimodal data fusion was developed to estimate lettuce phenotypic traits accurately
Reproduction assets foundThe paper's RGB-D lettuce images and trait measurements come from the publicly available Third Autonomous Greenhouse Challenge dataset deposited at 4TU.ResearchData, with an explicit availability statement and URL matching an allowed URL. No author analysis code or trained model is disclosed.Dataset · publicThis study used the Third Autonomous Greenhouse Challenge: Online Challenge Lettuce Images dataset publicly available at 4TU.ResearchData [ 36 ].Open asset ↗4TU.ResearchDatalines:819-832Plant phenotyping relevance match · UnverifiedbioRxiv · checked 15 Sept 2026
Due to their sessile nature, plants are unable to escape environmental factors that negatively impact health, resulting in losses to agricultural productivity. Rapid, non-invasive tools to detect plant stress response are essential for optimizing resource efficiency and mitigating the effects of extreme environmental pressures. However, many existing methods are either invasive, incompatible with other measurement techniques, or have not been applied to a wide range of varying environmental factors. In this study, we assess the physiological responses of four week old camelina (Camelina sativa) and sorghum (Sorghum bicolor) to chitosan, cold, drought, and both acute and chronic salt stress. Several plant characteristics were measured in parallel during stress exposure, including fluorescence and gas exchange parameters (MultispeQ and LI-6800), tissue electrical impedance with wearable biosensors (Multi-PIP), and biochemical properties via Fourier-transform infrared (FTIR) spectroscopy. We compiled unique profiles for whole plant physiological changes in response to environmental stress, demonstrating that certain aspects of plant health and makeup underwent alterations on differing temporal scales. This finding emphasizes the need for a comprehensive multi-modal approach to rapidly and accurately perform remote sensing of plant health in the field. Physiological parameters such as leaf impedance were also observed to rapidly change in response to treatment and can be leveraged to detect very early signs of plant perturbation. This research establishes the utility of a holistic phenotyping approach to inform agricultural strategies aimed at enhancing crop resilience under changing environmental conditions.
Why it matches plant phenotyping methods複数のセンサー・分光法を統合した非侵襲的な植物ストレス表現型取得と、マルチモーダル表現型解析の有用性が研究の中心である。
abstractRapid, non-invasive tools to detect plant stress response are essential
Automatically identifying key physiological factors in plants, such as leaf relative humidity (LRH), chlorophyll content (Chl), and nitrogen levels (N), is vital for effective aeroponic management and improving growth, yield, quality, and sustainability. Meta-learning (MetaL) solutions utilize data fusion and intelligent processing, ensuring fast and consistent outcomes. This paper aims to develop a novel MetaL framework that leverages multimodal data sources—including spectral, thermal, and IoT environmental data—to enable real-time, non-invasive identification of LRH, Chl, and N content in aeroponically grown lettuce. The research examined various spectral reflectance indices (SRIs) and thermal indicators from plant characteristics. Model-based feature selection was implemented using back-propagation neural networks (BPNN), decision trees (DT), and gradient boosting machines (GBM) to identify key attributes and optimize hyperparameters. The experimental findings indicated that deploying GBM-based top variables as the foundational model, combined with BPNN as the meta-model, significantly improved the accuracy of analyzing the assigned factors. The prediction scores (R²) for LRH, Chl, and N increased to 0.875 (RMSE=0.879), 0.886 (RMSE=0.694), and 0.930 (RMSE=0.184), respectively, compared to applying BPNN-based features alone as a standalone model. Overall, the designed methodology contributes to more accurate predictions of plant physiological states, enabling proactive steps toward sustainable aeroponic agriculture.
Why it matches plant phenotyping methodsスペクトル・熱画像・IoTデータを融合し、レタスの生理状態を非破壊推定するMetaL手法の開発が中心であり、植物表現型の取得・推定方法に該当する。
abstractThis paper aims to develop a novel MetaL framework that leverages multimodal data sources—including spectral, thermal, and IoT environmental data—to enable real-time, non-invasive identification of LRH, Chl, and N content in aeroponically grown lettuce.
The identification and control of plant diseases have become increasingly complex as agricultural lands expand globally. Traditional manual methods of monitoring and diagnosing diseases are not feasible for large-scale farms. The complexity is further heightened by the heterogeneity of data types collected through remote sensing methods, such as images, videos, and sensor readings, which need to be processed in real-time. This paper introduces a novel EdgeCloud Remote Sensing architecture that utilizes Deep Neural Networks (DNNs) with Transfer Learning to enhance the detection of plant diseases using remote sensing data. The proposed system employs a Fuzzy Deep Convolutional Neural Network (FCDCNN) to process multimodal data collected from both edge nodes and satellite point clouds. Transfer learning allows the sharing of model weights between cloud and edge nodes, thus optimizing performance without requiring frequent retraining. Simulations show that the proposed system achieves a 98% detection accuracy and reduces processing time by 25% compared to existing methods. By distributing tasks between edge and cloud resources, the system is able to process large datasets effectively and improve disease detection performance on vast farmlands.
Why it matches plant phenotyping methods植物病害状態をリモートセンシング画像・動画・センサーデータから推定する深層学習およびエッジ・クラウド基盤が研究の中心であり、病害フェノタイピング手法として採用する。
abstractThis paper introduces a novel EdgeCloud Remote Sensing architecture that utilizes Deep Neural Networks (DNNs) with Transfer Learning to enhance the detection of plant diseases using remote sensing data.
Field / plotLaboratory / benchtopMultimodalMultispectral / hyperspectralThermalWhole plant / canopy / plot / field
Multispectral imaging (MSI) is a technique used to inspect materials properties in different domains, ranging from industrial to medical and cultural heritage and, recently, precision agriculture. Even though several MSI solutions are already commercially available, the research community is working to optimize multispectral cameras in terms of performance and cost. Systems for the agricultural field are usually very compact, combined with drones for large areas acquisition. In this work, we detail the implementation of an innovative, modular and low-cost solution of a multispectral camera based on three core camera systems in the optical (VIS-NIR) and thermal (LWIR) range. Multispectral imaging is performed with a rotating wheel of interchangeable band-pass filters. The system is also equipped with a set of environmental sensors to acquire CO 2 concentration values, light intensity, temperature, and relative humidity of the surrounding environment. The technology and the measurement protocol were experimentally validated in laboratory and in open field. Advantages with respect to the available MSI cameras mounted on UAV is the integrated imaging in both the reflectance and the thermal emissive band in a close-up imaging setup and the use of environmental sensors. From the multispectral stack the spectral signature of the plants can be obtained and various vegetation indices (e.g., NDVI, NDRE) can be calculated for investigating the health status of the plant, while thermography provide additional monitoring. Close-up multispectral imaging is expected to tackle the new challenges of precision agriculture by enabling the acquisition of high-quality dataset on single plants.
Why it matches plant phenotyping methods植物のマルチスペクトル・熱画像カメラと環境センサーを開発し、測定プロトコルを実験的に検証して、植物のスペクトル特性・植生指数・健康状態を取得する方法が中心である。
abstractIn this work, we detail the implementation of an innovative, modular and low-cost solution of a multispectral camera based on three core camera systems in the optical (VIS-NIR) and thermal (LWIR) range.
Abstract Context The invasion of annual grasses in western U.S. rangelands promotes high litter accumulation throughout the landscape that perpetuates a grass-fire cycle threatening biodiversity. Objectives To provide novel evidence on the potential of fine spatial and structural resolution remote sensing data derived from Unmanned Aerial Vehicles (UAVs) to separately estimate the biomass of vegetation and litter fractions in sagebrush ecosystems. Methods We calculated several plot-level metrics with ecological relevance and representative of the biomass fraction distribution by strata from UAV Light Detection and Ranging (LiDAR) and Structure-from-Motion (SfM) datasets and regressed those predictors against vegetation, litter, and total biomass fractions harvested in the field. We also tested a hybrid approach in which we used digital terrain models (DTMs) computed from UAV LiDAR data to height-normalize SfM-derived point clouds (UAV SfM-LiDAR). Results The metrics derived from UAV LiDAR data had the highest predictive ability in terms of total (R2 = 0.74) and litter (R2 = 0.59) biomass, while those from the UAV SfM-LiDAR provided the highest predictive performance for vegetation biomass (R2 = 0.77 versus R2 = 0.72 for UAV LiDAR). In turn, SfM and SfM-LiDAR point clouds indicated a pronounced decrease in the estimation performance of litter and total biomass. Conclusions Our results demonstrate that high-density UAV LiDAR datasets are essential for consistently estimating all biomass fractions through more accurate characterization of (i) the vertical structure of the plant community beneath top-of-canopy surface and (ii) the terrain microtopography through thick and dense litter layers than achieved with SfM-derived products.
Why it matches plant phenotyping methodsUAV LiDARおよびSfMから植生・リター・総バイオマスを推定する測定手法を開発・比較評価しており、植物形質取得が研究の中心である。
abstractTo provide novel evidence on the potential of fine spatial and structural resolution remote sensing data derived from Unmanned Aerial Vehicles (UAVs) to separately estimate the biomass of vegetation and litter fractions in sagebrush ecosystems.
Preharvest crop yield estimation is crucial for achieving food security and managing crop growth. Unmanned aerial vehicles (UAVs) can quickly and accurately acquire field crop growth data and are important mediums for collecting agricultural remote sensing data. With the rapid development of machine learning, especially deep learning, research on yield estimation based on UAV remote sensing data and machine learning has achieved excellent results. This paper systematically reviews the current research of yield estimation research based on UAV remote sensing and machine learning through a search of 76 articles, covering aspects such as the grain crops studied, research questions, data collection, feature selection, optimal yield estimation models, and optimal growth periods for yield estimation. Through visual and narrative analysis, the conclusion covers all the proposed research questions. Wheat, corn, rice, and soybeans are the main research objects, and the mechanisms of nitrogen fertilizer application, irrigation, crop variety diversity, and gene diversity have received widespread attention. In the modeling process, feature selection is the key to improving the robustness and accuracy of the model. Whether based on single modal features or multimodal features for yield estimation research, multispectral images are the main source of feature information. The optimal yield estimation model may vary depending on the selected features and the period of data collection, but random forest and convolutional neural networks still perform the best in most cases. Finally, this study delves into the challenges currently faced in terms of data volume, feature selection and optimization, determining the optimal growth period, algorithm selection and application, and the limitations of UAVs. Further research is needed in areas such as data augmentation, feature engineering, algorithm improvement, and real-time yield estimation in the future.
Why it matches plant phenotyping methodsUAVリモートセンシングと機械学習による作物収量推定手法を体系的にレビューしており、植物形質の取得・推定方法が中心である。
abstractThis paper systematically reviews the current research of yield estimation research based on UAV remote sensing and machine learning
In the context of climate change, extreme weather events, represented by frost injury, are increasingly having a negative impact on the growth of tea plants. This has brought huge losses to the tea industry. Traditionally, the freezing injury of tea plants in the field was assessed by vision. This is labor-intensive and subjective. In this research, multimodal remote sensing data from different periods of natural overwintering tea plantations were collected by using unmanned aerial vehicles (UAV) equipped with multispectral (MS), thermal infrared (TIR) and RGB sensors. And the physiological data of tea leaves on the same day were obtained to construct a tea cold injury score (TCIS). Then, a convolutional neural networks-gate recurrent unit (CNN-GRU) model was improved for estimating TCIS. To better compare the performance of CNN-GRU, a single GRU model and three classical machine learning models were also used for comparison. The study found that: (1) The multimodal data fusion was superior to the unimodal data. The best prediction results were achieved for the combined bimodal MS + RGB data (Rp² = 0.862, RMSEP = 0.138, RPD = 2.220); (2) The CNN-GRU hybrid model was superior to the other four baseline models. The best effect was achieved based on the multivariate input of MS + RGB (Rp² = 0.862) or MS + RGB + TIR (Rp² = 0.850); (3) The accuracy of the model after removing soil features was lower than that of the model without background removal. Therefore, the TCIS-CNN-GRU model combined with multi-source remote sensing data can objectively and accurately evaluate the cold injury phenotype of tea plants, making the CNN-GRU model more scientific and promising.
Why it matches plant phenotyping methodsUAVマルチセンサー画像から茶樹の凍害表現型を推定する取得・計算手法を開発し、複数モデルと比較検証しており、フェノタイピング手法が中心である。
abstractmultimodal remote sensing data from different periods of natural overwintering tea plantations were collected by using unmanned aerial vehicles (UAV) equipped with multispectral (MS), thermal infrared (TIR) and RGB sensors.
The tomato plant’s main-stem is a feasible lead for robotic searching the grows discretely-growing targets of harvesting, pruning or pollinating. Owing to the highlighted reflection characteristics of the main-stem in the near-infrared (NIR) waveband, this study proposes a multimodal hierarchical fusion method (YOLACTFusion) based on the attention mechanism, to achieve an instance segmentation of the main-stem from similar-colored differentiation (i.e., green leaf and green fruit) in robotic vision systems. The model inputs RGB images and 900–1100 nm NIR images into two ResNet50 backbone networks and uses a parallel attention mechanism to fuse feature maps of various scales together into the head network, to improve the segmentation performance of the main-stem of RGB images. The loss function for the multimodal image weights the original loss on the RGB image and the position offset loss and classification loss on the NIR image. Furthermore, the local depthwise separable convolution is used for the backbone network, and Conv-BN layers are merged to reduce the computational complexity. The results show that the precision and recall of YOLACTFusion of the main-stem detection, respectively reached 93.90 % and 62.60 %; and the precision and recall of instance segmentation reached 95.12 % and 63.41 %, respectively. Compared to YOLACT, the mean average precision (mAP) of YOLACTFusion is increased from 39.20 % to 46.29 %, the model size is reduced from 199.03 MB to 165.52 MB, while the image processing efficiency remains similar. The overall results show that the multimodal instance segmentation method proposed in this study significantly improves the detection and segmentation of tomato main-stems under a similar-colored background, which would be a potential method for improving agricultural robot’s visual perception.
Why it matches plant phenotyping methodsRGB-NIR画像を用いてトマト主茎をセグメンテーションする手法を開発・評価しており、植物器官の形態取得が中心的な技術貢献である。
abstractthis study proposes a multimodal hierarchical fusion method (YOLACTFusion) based on the attention mechanism, to achieve an instance segmentation of the main-stem
Quantitative remote sensing of crop diseases at the field or plot scale is essential for crop management. Conventional approaches frequently rely solely on single-modal remote sensing data, resulting in performance limitations. Investigating the utilization of multi-modal imagery, including high spatial resolution Red–Green–Blue (RGB) images, high spectral resolution multispectral (MS) images, and vegetation indexes (VIs) informed by expert knowledge, to enhance quantitative inversion warrants further research. In this study, we propose a novel DL-based quantitative framework for analyzing multimodal remote sensing images, named RustQNet. This framework facilitates precise and high-throughput quantitative assessment of wheat stripe rust (WSR) disease index (DI). RustQNet is designed to handle multimodal remote sensing images, such as RGB, MS, and VIs . To improve the fusion of different modalities, RustQNet incorporates a mutual information minimization (MIM) module, which encourages information complementarity across modalities in a more compact fashion. To better train our model, we construct a benchmark dataset of time-series multimodal UAV images covering the entire period of the WSR epidemic, from the initial infection to the severe outbreaks. This dataset contains over 180 million quantitatively annotated pixels, making it the most comprehensive dataset currently available for the quantitative inversion of WSR. The results show that the R2 value of the RGB+MS+VI three-modal model is 0.8024, which is improved 17.65%–35.59% compared with the single-modality model, and 1.27%–6.67% compared with the bimodal model. The RMSE of the three-modal model is 11.4776, which is 20.51%–30.71% lower than the single-modality model and 3.93%–11.56% lower than the bimodal model. When DI<=10, the RMSE is 5.47, indicating potential for early detection. In comparison to typical semantic segmentation models and sophisticated multimodal algorithms, the RustQNet exhibited exceptional performance. The findings suggest that RustQNet may serve as a multimodal data processing tools for precise and efficient crop disease quantitative inversion, thereby offering valuable insights for the high-throughput quantitative inversion of additional phenotypic traits, including yield and plant height.
Why it matches plant phenotyping methods植物体の病害状態(コムギ赤さび病の病害指数)をUAVマルチモーダル画像から定量推定する手法を開発し、ベンチマークデータセットで評価しているため、植物フェノタイピング手法が中心である。
abstractwe propose a novel DL-based quantitative framework for analyzing multimodal remote sensing images, named RustQNet.
Real-time monitoring of leaf area index (LAI) in cotton (Gossypium hirsutum L.) plays a vital role in guiding field fertilization, water management, growth observation and yield prediction. The use of unmanned aerial vehicles (UAVs) equipped with diverse sensors enables flexible and rapid LAI measurement across extensive areas. This study evaluates the efficacy of UAV-mounted LiDAR and high-resolution camera-derived point cloud data in LAI prediction. We assessed the integration of canopy spectral-textural characteristics from multispectral data with structural features from point cloud data for LAI forecasting. Furthermore, we compared various machine learning and deep learning models, selected the optimal one, and applied the SHAP (Shapley Additive Explanations) method to identify key features and their influence patterns in this model. The findings are distilled into four key points: (1) The performance of canopy structure metrics based on the two sensors varied across fertility periods due to differences in canopy closure; (2) The DNN-based (Deep Neural Network) LAI prediction model excelled with a single-period dataset, achieving an R² of 0.81 and an RMSE% of 11.36 %. Similarly, in full-period multimodal data fusion, it demonstrated superior performance, evidenced by an R² of 0.84 and an RMSE% of 9.94 %. (3) Compared to the unimodal data model, the multimodal data model yielded superior results and exhibited greater robustness. (4) In the DNN-based LAI prediction model utilizing multimodal data, texture features contributed most significantly. The results suggest that the DNN model, when employing multimodal data fusion, offers not only relatively precise and robust estimates of crop LAI but also contributes valuable insights for crop phenotyping and enhanced field management. This approach subsequently improves spatial prediction accuracy and the quality of decision-making in crop production.
Why it matches plant phenotyping methodsUAV LiDAR・マルチスペクトル画像・点群を統合し、機械学習で綿花のLAIという植物形質を推定する手法を比較・評価しており、フェノタイピング手法が中心である。
abstractThis study evaluates the efficacy of UAV-mounted LiDAR and high-resolution camera-derived point cloud data in LAI prediction.
INTRODUCTION: Detecting and monitoring crop stress is crucial for ensuring sufficient and sustainable crop production. Recent advancements in unoccupied aerial vehicle (UAV) technology provide a promising approach to map key crop traits indicative of stress. While using single optical sensors mounted on UAVs could be sufficient to monitor crop status in a general sense, implementing multiple sensors that cover various spectral optical domains allow for a more precise characterization of the interactions between crops and biotic or abiotic stressors. Given the novelty of synergistic sensor technology for crop stress detection, standardized procedures outlining their optimal use are currently lacking. MATERIALS AND METHODS: This study explores the key aspects of acquiring high-quality multi-sensor data, including the importance of mission planning, sensor characteristics, and ancillary data. It also details essential data pre-processing steps like atmospheric correction and highlights best practices for data fusion and quality control. RESULTS: Successful multi-sensor data acquisition depends on optimal timing, appropriate sensor calibration, and the use of ancillary data such as ground control points and weather station information. When fusing different sensor data it should be conducted at the level of physical units, with quality flags used to exclude unstable or biased measurements. The paper highlights the importance of using checklists, considering illumination conditions and conducting test flights for the detection of potential pitfalls. CONCLUSION: Multi-sensor campaigns require careful planning not to jeopardise the success of the campaigns. This paper provides practical information on how to combine different UAV-mounted optical sensors and discuss the proven scientific practices for image data acquisition and post-processing in the context of crop stress monitoring.
Why it matches plant phenotyping methods作物ストレスという植物状態を推定するUAVマルチセンサーの取得、校正、前処理、データ融合、品質管理の実践的方法を主題としており、フェノタイピング手法が中心である。
abstractThis paper provides practical information on how to combine different UAV-mounted optical sensors and discuss the proven scientific practices for image data acquisition and post-processing in the context of crop stress monitoring.
AppleField / plotMultimodalLiDAR / point cloudRGB / grayscaleFruitCounting2D/3D reconstructionTrackingGrowth / development / phenology
Monitoring orchards at the individual tree or fruit level throughout the growth season is crucial for plant phenotyping and horticultural resource optimization, such as chemical use and yield estimation. We present a 4D spatio-temporal metric-semantic mapping system that integrates multi-session measurements to track fruit growth over time. Our approach combines a LiDAR-RGB fusion module for 3D fruit localization with a 4D fruit association method leveraging positional, visual, and topology information for improved data association precision. Evaluated on real orchard data, our method achieves a 96.9% fruit counting accuracy for 1,790 apples across 60 trees, a mean fruit size estimation error of 1.1 cm, and a 23.7% improvement in 4D data association precision over baselines. We publicly release a multimodal dataset covering five fruit species across their growth seasons at https://4d-metric-semantic-mapping.org/
Why it matches plant phenotyping methods果実の成長追跡・計数・サイズ推定という植物器官形質を対象に、LiDAR-RGB融合と4D対応付け手法を開発・評価し、データセットも公開しているため、フェノタイピング手法が中心です。
abstractWe present a 4D spatio-temporal metric-semantic mapping system that integrates multi-session measurements to track fruit growth over time.
Growth monitoring of crops is a crucial aspect of precision agriculture, essential for optimal yield prediction and resource allocation. Traditional crop growth monitoring methods are labor-intensive and prone to errors. This study introduces an automated segmentation pipeline utilizing multi-date aerial images and ortho-mosaics to monitor the growth of cauliflower crops (Brassica Oleracea var. Botrytis) using an object-based image analysis approach. The methodology employs YOLOv8, a Grounding Detection Transformer with Improved Denoising Anchor Boxes (DINO), and the Segment Anything Model (SAM) for automatic annotation and segmentation. The YOLOv8 model was trained using aerial image datasets, which then facilitated the training of the Grounded Segment Anything Model framework. This approach generated automatic annotations and segmentation masks, classifying crop rows for temporal monitoring and growth estimation. The study’s findings utilized a multi-modal monitoring approach to highlight the efficiency of this automated system in providing accurate crop growth analysis, promoting informed decision-making in crop management and sustainable agricultural practices. The results indicate consistent and comparable growth patterns between aerial images and ortho-mosaics, with significant periods of rapid expansion and minor fluctuations over time. The results also indicated a correlation between the time and method of observation which paves a future possibility of integration of such techniques aimed at increasing the accuracy in crop growth monitoring based on automatically derived temporal crop row segmentation masks.
Why it matches plant phenotyping methods航空画像・オルソモザイクから作物列を自動セグメンテーションし、時系列の生育・成長を推定する画像解析パイプラインが研究の中心であり、植物形質の取得手法として適格。
abstractThis study introduces an automated segmentation pipeline utilizing multi-date aerial images and ortho-mosaics to monitor the growth of cauliflower crops
Reproduction assets foundThe paper's Data Availability Statement points to the authors' public Mendeley Data repository (GobhiSet, DOI 10.17632/dcjjcwc5dh.4), which contains the raw, manually, and automatically annotated RGB aerial images and ortho-mosaics of cauliflower used for the YOLOv8x-seg and Grounded SAM training and growth analysis inDataset · publicon of the manuscript.
Funding: This research received no external funding.
Data Availability Statement: No new data was created. However, the data that were used to perform
this research can be found in the article published at https://doi.org/10.1016/j.dib.2024.110506 and
available in the repository DOI: 10.17632/dcjjcwc5dh.4 (https://data.mendeley.com/drafts/dcjjcwc5dh).Conflicts of Interest: The authors declare no conflicts of interest.
References
1. Di, L.; Ustundag, B. Crop Growth Modeling and Yield Forecasting. In Agro-Geoinformatics; Springer: Cham, Switzerland, 2021.
[CrossRef]
2. Mithen, S.; Jenkins, E.; Jamjoum, K.; Nuimat, S.; Nortcliff, S.; Finlayson, B. Experimental crop growingOpen asset ↗data.mendeley.com · 10.17632/dcjjcwc5dh.4pdf-raw-page:17 lines:1-52Plant phenotyping relevance match · UnverifiedEurope PMC · checked 7 Sept 2026
Background Wheat (Triticum aestivum L.) is an important grain crops in the world, and its growth and development in different stages is seriously affected by saline-alkali stress, especially in seedling stage. Therefore, nondestructive detection of wheat seedlings under saline-alkali stress can provide more comprehensive technical support for wheat breeding, cultivation and management. Results This research focused on moisture signal prediction and classification of saline-alkali stress in wheat seedlings using fusion techniques. After collecting and analyzing transverse relaxation time and Multispectral imaging (MSI) information of wheat seedlings, four regression models were used to predict the moisture signal. K-Nearest Neighbor (KNN) and Gaussian-Naïve Bayes (GNB) models were combined with fivefold cross validation to classify the prediction of wheat seedling stress. The results showed that wheat seedlings would increase the bound water content through a certain mechanism to enhance their saline-alkali stress. Under the same Na concentration, the effect of alkali stress on moisture, growth and spectrum of wheat seedlings is stronger than salt stress. The Gradient Boosting Decision Regression Tree model performs the best in predicting wheat moisture signals, with a coefficient of determination (R2P) of 0.98 and a root mean square error of 109.60. It also had a short training time (1.48 s) and an efficient prediction speed (1300 obs/s). The KNN and GNB demonstrated significantly enhanced predictive performance when classifying the fused dataset, compared to using single datasets individually. In particular, the GNB model performing best on the fused dataset, with Precision, Recall, Accuracy, and F1-score of 90.30, 88.89%, 88.90%, and 0.90, respectively. Conclusions Under the same Na concentration, the effects of alkali stress on water content, spectrum, and growth of wheat were stronger than that of salt stress, which was more unfavorable to the growth of wheat. The fusion of low-field nuclear magnetic resonance and MSI technology can improve the classification of wheat stress, and provide an effective technical method for rapid and accurate monitoring of wheat seedlings under saline-alkali stress.
Why it matches plant phenotyping methods低磁場NMRとマルチスペクトル画像を融合し、コムギ幼植物の水分状態と塩・アルカリストレスを非破壊推定・分類する手法が研究の中心であるため。
abstractThis research focused on moisture signal prediction and classification of saline-alkali stress in wheat seedlings using fusion techniques.