← PhenoCode Atlas

Unverified paper discovery

Plant phenotyping methods.

植物形質を測っただけの研究ではなく、フェノタイピング手法の開発・検証・実質的利用・ベンチマーク・方法レビューとの関連性が見つかった論文を中心に表示します。

表示条件: Multimodal条件を解除 ×
53 papers · code / dataset availability confirmedLatest completed run · 2016-01-01 – 2026-09-13

自動判定された未検証候補です。Catalogへの掲載にはキュレーター承認が必要です。

Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published11 Sept 2026Plant physiology

Characterization of Rhizosphere Oxidation Associated with Root Development in Rice Using Planar Oxygen Optodes.

RiceMultimodalX-ray / CTRootMorphology / geometry measurementPhysiological trait estimationGrowth / time-series analysisGrowth / development / phenologyRoot system architecture

Rhizosphere oxidation is a key adaptive mechanism in reductive soil environments, in which oxygen released from roots alters rhizosphere redox conditions and regulates biogeochemical processes. Rice plants possess an internal oxygen transport system, and radial oxygen loss (ROL) from roots is closely associated with root development. However, the spatial patterns of ROL in soil and their relationships with root traits remain poorly characterized. In this study, we developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice. Daily time-course tracking of individual crown roots revealed dynamic changes in the spatial distribution and magnitude of rhizosphere oxygen in relation to root elongation and aging. Root thickness was positively correlated with dissolved oxygen levels near root tips. Genotypic comparisons further identified a cultivar with reduced rhizosphere oxidation despite possessing thicker roots among the tested genotypes, thereby indicating the involvement of additional physiological processes. Overall, these findings demonstrate that rhizosphere oxidation is regulated by root growth stage and thickness and dynamically modulated during root development.

Why it matches plant phenotyping methods平面酸素オプトードとX線CTを統合したマルチモーダル画像システムを開発し、根の発達と根圏酸化を時系列・個体別に定量化しており、表現型取得手法が研究の中心である。

abstractwe developed a multimodal imaging system that integrates planar oxygen optodes with X-ray computed tomography to simultaneously visualize rhizosphere oxidation and root development in rice.
Reproduction assets foundThe paper's Data availability statement explicitly deposits the authors' RG2DO-Root analysis program together with sample optode and CT images (the paper's phenotyping inputs) in a public GitHub repository, matching the allowed URL.
Code · publicing 8 This work was supported by project JPNP18016, commissioned by the New Energy and 9 Industrial Technology Development Organization (NEDO), JST CREST (JPMJCR17O1), 10 and JST ALCA-Next (JPMJAN23D3). 11 12 Data availability 13 The source code and sample data (optode and CT images) are available from the 14 GitHub repository (https://github.com/tsubasa-kawai28/RG2DO-Root).15 16 References 17 Aguilar EA et al. 2003. Oxygen distribution and movement, respiration and nutrient 18 loading in banana roots (Musa spp. L.) subjected to aerated and oxygen-depleted 19 environments. Plant Soil. 253:91–102. https://doi.org/10.1023/A:1024598319404.20 Armstrong W, Wright EJ. 1975. Radial oxygen loss fromOpen asset ↗https://github.com/tsubasa-kawai28/RG2DO-Root · RG2DO-Rootpdf-raw-page:19 lines:1-82
Code / dataset availability confirmedEurope PMC · Crossref · checked 15 Sept 2026
Published3 Sept 2026Methods in ecology and evolution

Mind(the)Plant: An expandable multimodal facility for the integrated characterization of plant behaviour

Growth chamberMultimodalStereoRootStem / branchTrackingGrowth / development / phenology

Understanding plant behaviour requires the integration of multiple phenotypic and physiological signals measured over time under controlled conditions. However, different plant signals are typically studied using separate experimental setups, limiting temporal alignment and integrative analyses.We present Mind(the)Plant, a modular experimental facility designed for the synchronized, long-term acquisition of multimodal plant data, including three-dimensional shoot kinematics, above- and below-ground volatile organic compounds (VOCs) and root imaging. Its modular architecture is designed to accommodate additional acquisition modules, such as electrophysiological signalling, as future extensions. The platform integrates a controlled growth environment with stereovision imaging, high-resolution time-of-flight mass spectrometry and custom rhizocameras. These components are connected through a unified network infrastructure that ensures synchronized acquisition and centralized data handling.We validate the performance of each acquisition module through multi-week recordings, demonstrating high-temporal stability, reliable stereovision synchronization, effective isolation of VOCs signals and robust operation of below-ground imaging. We further illustrate the analytical potential of the platform using a one-day continuous multimodal acquisition combining shoot kinematics, above-ground VOC emissions, rhizocameras observations and environmental data.Mind(the)Plant provides a novel methodological framework for studying plant behaviour, signalling and phenotypic plasticity in ecological and evolutionary research. By enabling coordinated measurements of multiple plant response modalities, the platform supports investigations of dynamic plant-environment and plant-plant interactions from a behavioural perspective.

Why it matches plant phenotyping methods植物の複数の表現型・生理シグナルを同期取得する施設を開発し、各取得モジュールの性能を検証しているため、表現型計測プラットフォームが研究の中心です。

abstractWe present Mind(the)Plant, a modular experimental facility designed for the synchronized, long-term acquisition of multimodal plant data, including three-dimensional shoot kinematics, above- and below-ground volatile organic compounds (VOCs) and root imaging.
Reproduction assets foundThe paper's data availability statement explicitly deposits data, code and processing pipelines (supporting the multimodal plant phenotyping measurements and analysis) in a public Zenodo archive with an authors' URL matching an allowed URL.
Code · publicf Interest Statement The authors have no conflicts of interest to declare. Peer Review The peer review history for this article is available at https://www.webofscience.com/api/gateway/wos/peer-review/10.1111/2041-210x.70411 . Data availability Statement Data, code and processing pipelines supporting this study are available at https://doi.org/10.5281/zenodo.22095454 ( Simonetti & Castiello, 2026 ). References Avesani S, Bonato B, Simonetti V, Guerra S, Ravazzolo L, Gjinaj G, Dadda M, Castiello U. Comparing proton transfer reaction (PTR) and adduct ionization mechanism (AIM) for the study of volatile organic compounds. Molecules. 2026;31(3):402. doi: 10.3390/molecules31030402. Baluška F, LeOpen asset ↗zenodo · 10.5281/zenodo.22095454lines:482-508
Code / dataset availability confirmedOpenAlex · checked 5 Sept 2026
Published18 Aug 2026Journal of King Saud University - Computer and Information SciencesCited by 0 · OpenAlex ↗

A residual forecasting framework for plant dynamic growth based on cross-modal spatial alignment

MaizeWheatField / plotMultimodalWhole plant / canopy / plot / fieldGrowth / time-series analysisGrowth / development / phenologyPlant / canopy height

Plant phenotyping is essential for modern crop breeding, yet traditional static image analysis fails to capture the nonlinear dynamics of plant growth. Existing time-series forecasting models exhibit notable limitations when processing multimodal data: global pooling operations may compress local 2D spatial topology of plants, and shallow feature concatenation may be insufficient for effective cross-modal semantic alignment. Moreover, current methods typically regress absolute morphological states, which may contribute to temporal lag during nonlinear growth spurts. In this paper, we propose ST-CrossGro-Former, a cross-modal residual forecasting framework for plant dynamic growth. The network removes the final global pooling and classification layers to preserve spatial topology and incorporates a scalar-guided cross-modal attention module based on the standard query-key-value formulation. This module utilizes 1D morphological features as queries to dynamically weight local visual regions, promoting multimodal feature alignment. Concurrently, a residual incremental forecasting strategy is introduced to predict short-term growth increments rather than absolute states, aiming to improve tracking sensitivity to sudden growth events. Evaluations on the UNL-CPPD maize dataset and supplementary validation on the FIP1 wheat field dataset show that the proposed model achieves competitive single-step forecasting accuracy and favorable temporal trajectory alignment compared with adapted spatiotemporal attention, graph-based, and physics-informed baselines under the evaluated settings. In particular, the FIP1 results suggest that ST-CrossGro-Former can maintain favorable height trajectory alignment under a field-acquired wheat setting, indicating its potential for helping mitigate temporal misalignment in dynamic growth forecasting.

Why it matches plant phenotyping methods植物の動的形態成長を予測する新規クロスモーダル手法を開発し、トウモロコシ・コムギデータセットで評価しており、表現型の抽出・予測手法が中心である。

abstractwe propose ST-CrossGro-Former, a cross-modal residual forecasting framework for plant dynamic growth.
Reproduction assets foundThe paper evaluates its ST-CrossGro-Former model on two public plant phenotyping datasets: the UNL-CPPD maize dataset (explicitly stated as publicly available with a repository URL) and the FIP1 wheat field dataset (public dataset from ETH Zürich, with its GigaScience dataset publication DOI). No author analysis code,
Dataset · publicThe UNL-CPPD dataset used in this research was acquired from the UNL Plant Phenotyping Datasets repository, accessible at https://plantvision.unl.edu/datasets.Open asset ↗UNL Plant Phenotyping Datasets · UNL-CPPDlines:266-273
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published13 Aug 2026Frontiers in Plant ScienceCited by 0 · OpenAlex ↗

AI driven multi modal deep learning system for wheat disease detection, yield prediction, and crop health monitoring

WheatField / plotGreenhouseMultimodalPanicle / ear / spikeWhole plant / canopy / plot / fieldClassificationCountingObject detectionStress / disease detection

Sustainable wheat farming is challenging. Real-time information on crop health, disease transmission, and anticipated yields is essential for farmers. However, they frequently use slow, expensive, or non-communicative tools. This project develops a workable solution. There is no need for massive server farms because the entire system operates on a single graphics card. It incorporates images of wheat fields, Indian farming notes, greenhouse records, harvest statistics, and NASA meteorological data. Consider them as various “eyes” for crop photo analysis, and we tried several lightweight computer vision models. ConvNeXt-Tiny was slower but could operate on older equipment with 75% accuracy; EfficientNetB0 recognised wheat heads with 92% accuracy; and AgroMark, a hybrid solution that merged photo analysis with agricultural metadata (soil type, rainfall, increased to 87%, etc. Combining picture analysis with attention mechanisms (CBAM) allowed us to anticipate the amount of wheat that a field will yield based on these photo insights, and the results showed that our predictions were accurate, with an R 2 score of 0.97. Additionally, we developed a versatile detector that simultaneously detects disease, stress, head count, and pests. It is adjusted to deal with training data that is unbalanced (some diseases are common, while others are rare). As we packed everything into a 16-GB graphics card, we spent real time determining which strategies smaller training sets, removing weak features, and adjusting loss functions, work. We encounter real-world obstacles along the road, such as photographs from different locations not always match, mislabeled photographs from different locations not always match, mislabeled diseases, and neglected rare pests. Our step-by-step instructions, charts, and code are available.

Why it matches plant phenotyping methods小麦画像から病害・ストレス・穂数・収量などの植物形質・状態を推定するマルチモーダル手法を開発し、複数モデルの精度比較と実装上の検証を行っており、表現型取得・推定が研究の中心である。

abstractThis project develops a workable solution.
Reproduction assets foundThe paper builds its multimodal wheat phenotyping analysis on several explicitly cited public data assets: the Kaggle Wheat Plant Diseases image dataset (used for disease classification, Tables 2 and 9), the Global Wheat Head Detection dataset (used for head detection, Tables 1 and 6), FAOSTAT and India Open Government
Dataset · publicAvailable online at: https://www.fao.org/faostat/ . FAOSTAT statistical database.Open asset ↗lines:1110-1162
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published27 Jul 2026Frontiers in Fungal BiologyCited by 0 · OpenAlex ↗

AgriFusionNet: a context-aware multimodal leaf disease diagnosis and classification system for sustainable plant health monitoring

Growth chamberMultimodalLeafClassificationStress / disease detectionDisease symptoms / severity

Early and accurate identification of plant diseases is essential for improving crop productivity and ensuring food security. Many existing deep learning-based plant disease classification methods rely solely on leaf images collected from a controlled environment, which limits their applicability in real-world agricultural conditions where symptoms may be visually unclear and influenced by environmental factors. To address these challenges, this study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification. The proposed architecture employs EfficientNet-B0 for visual feature extraction, BERT for semantic representation of symptom descriptions, and a lightweight multilayer perceptron for modeling environmental factors such as temperature, humidity, rainfall, and soil moisture. Features from all three modalities are fused into a unified representation to train the CNN model. The model is trained and tested upon the Context-Aware Multimodal Augmented PlantVillage dataset covering 38 plant diseases and healthy classes. Experimental results show that AgriFusionNet gives an overall accuracy of 98.94% on the dataset Context-Aware Multimodal Augmented PlantVillage, with competitive precision and recall and F1-score. The multimodal framework facilitates the co-learning of visual, semantic, and contextual environmental representations and the analyses of the confusion matrix and feature interactions give insights into cross-modal relationships. The proposed approach aims to explore context-aware multimodal representation learning for agricultural AI applications, with emphasis on integrating complementary visual, semantic, and contextual information.

Why it matches plant phenotyping methods葉画像を中心に、症状記述と環境情報を統合して植物病害状態を分類する手法を開発・評価しており、植物フェノタイピング手法が中心である。

abstractthis study discusses AgriFusionNet, a context-aware multimodal deep learning framework that integrates leaf images, textual symptom descriptions, and environmental data for robust plant disease classification.
Reproduction assets foundThe paper's data availability statement points to the Context-Aware Multimodal Augmented PlantVillage dataset (leaf images, symptom text, environmental data used for the phenotyping/classification analysis) deposited publicly on IEEE Dataport with a DOI matching an allowed URL.
Dataset · publicPublicly available datasets were analyzed in this study. This data can be found here: Dataset. IEEE Dataport. https://dx.doi.org/10.21227/9jat-r836 [Accessed on August 2025].Open asset ↗IEEE Dataport · 10.21227/9jat-r836lines:1029-1047
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published16 Jun 2026Frontiers in Computer ScienceCited by 0 · OpenAlex ↗

Hybrid multimodal learning framework for crop disease detection, adaptive treatment, and price forecasting

CottonTomatoMultimodalLeafClassificationObject detectionStress / disease detectionDisease symptoms / severity

Crop diseases play a significant role in food production globally; therefore, there is an urgent need to develop quick and accurate diagnostic techniques that are more effective than manual inspection methods. The proposed hybrid multimodal learning framework in this research provides a solution that integrates adaptive therapy suggestion, market price prediction, and image-based disease detection. This study also proposes a framework for pesticide recommendation and the treatment of plants. This study experiment on tomato and cotton crop leaf data for disease detection. Experimental results on a tomato crop disease detection dataset show that the proposed model shows high performance. EfficientNetB0 provides more stability and generalization capabilities in different scenarios compared to other models, such as YOLOv8, ResNet50, and a custom CNN model. The use of a knowledge-based decision support system provides sustainable pesticide recommendations based on environmental and symptom-specific parameters. Forecasting of pesticide prices through LSTM methods yields forecasts within 3.2% and 4.1% MAE, enabling improved decision-making by providing instant points of reference for potential price movements. Research uses SHAP and LIME to provide explainability to users, thus improving user buy-in through transparency. Overall, this modular system provides a data-driven decision-making model to improve the efficiency of managing crops.

Why it matches plant phenotyping methods植物葉画像から病害状態を推定する画像ベース手法を、複数モデルで比較評価しており、植物病害フェノタイピングがシステムの主要構成要素です。価格予測や農薬推薦も含みますが、病害検出の技術評価が明示されています。

abstractThe proposed hybrid multimodal learning framework in this research provides a solution that integrates adaptive therapy suggestion, market price prediction, and image-based disease detection.
Reproduction assets foundThe paper's disease-detection experiments use publicly available cotton and tomato leaf image datasets (Kaggle, IEEE DataPort, Roboflow), all cited with explicit public URLs in the references. No author analysis code or trained model checkpoints are stated as publicly available; the supplementary material is referenced
Dataset · publiccholar View reference in article 19 Muppala C. Guruviah V. ( 2020 ). Machine vision detection of pests, diseases, and weeds: a review . J. Phytol. 12 , 9 – 19 . doi: 10.25081/jp.2020.v12.6145 CrossRef Google Scholar View reference in article 20 National College of Ireland ( 2025 ). “Cotton Disease Dataset.” Available online at: https://www.kaggle.com/datasets/janmejaybhoi/cotton-disease-dataset (Accessed May 19, 2025). Google Scholar View reference in article 21 Naveed Gul and Kaggle ( 2026 ). Tomato Leaf Disease . Kaggle. Available online at: https://www.kaggle.com/datasets/naveedgull/tomato-leaf-disease (Accessed March 29, 2026). Google Scholar View reference in article 22 Ngugi H. N. EzugOpen asset ↗Kagglelines:554-633
Dataset · publicreference in article 20 National College of Ireland ( 2025 ). “Cotton Disease Dataset.” Available online at: https://www.kaggle.com/datasets/janmejaybhoi/cotton-disease-dataset (Accessed May 19, 2025). Google Scholar View reference in article 21 Naveed Gul and Kaggle ( 2026 ). Tomato Leaf Disease . Kaggle. Available online at: https://www.kaggle.com/datasets/naveedgull/tomato-leaf-disease (Accessed March 29, 2026). Google Scholar View reference in article 22 Ngugi H. N. Ezugwu A. E. Akinyelu A. A. Abualigah L. ( 2024 ). Revolutionizing crop disease detection with computational deep learning: a comprehensive review . Environ. Monit. Assess. 196 : 302 . doi: 10.1007/s10661-024-12454-z Pubmed AOpen asset ↗Kagglelines:554-633
Dataset · publicComputer Vision and Pattern Recognition (CVPR) ( Las Vegas, NV : IEEE ), 779 – 788 . doi: 10.1109/CVPR.2016.91 CrossRef Google Scholar View reference in article 29 Roboflow ( 2026a ). A Comprehensive Dataset of Cotton Plant Diseases for National Disease Identification and Treatment Guidance | IEEE DataPort. Available online at: https://ieee-dataport.org/documents/comprehensive-dataset-cotton-plant-diseases-national-disease-identification-and-treatment (Accessed March 29, 2026). Google Scholar View reference in article 30 Roboflow ( 2026b ). Cotton Plant Disease Prediction Object Detection Model by National College of Ireland . Available online at: https://universe.roboflow.com/national-colleOpen asset ↗IEEE DataPortlines:554-633
Code / dataset availability confirmedEurope PMC · checked 8 Sept 2026
Published21 May 2026Scientific dataCited by 0 · OpenAlex ↗

A Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis.

Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationCountingGrowth / development / phenology

Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.

Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画データセットであり、自動フェノロジー検出と空間検証のためのベンチマーク基盤が中心である。

abstractThe Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions.
Reproduction assets foundThe paper describes a public multi-modal Actinidia chinensis phenology dataset (annotated images, georeferenced videos, ground-truth counts) deposited on Zenodo, plus authors' MIT-licensed preprocessing scripts on GitHub. CVAT and FiftyOne are generic third-party tools and excluded.
Dataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available at Zenodo: https://doi.org/10.5281/zenodo.17371025.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:12 lines:1-92
Code · publicCustom scripts for dataset preparation are publicly available under the MIT License at https://github.com/Open asset ↗GitHubpdf-page:12 lines:1-92
Code / dataset availability confirmedCrossref · Europe PMC · checked 15 Sept 2026
Published21 May 2026Scientific ReportsCited by 0 · OpenAlex ↗

Hybrid deep learning-based multimodal framework for plant leaf disease classification using RGB, Excess Green (ExG), and pseudo-thermal representations with MobileNetV2

MultimodalRGB / grayscaleThermalLeafClassificationCalibration / preprocessingStress / disease detectionDisease symptoms / severity

Abstract Plant diseases are a serious danger to the world’s food security, because they lower agricultural output and increase economic losses. Due to subjectivity, fluctuating lighting, and environmental unpredictability, traditional visual examination techniques are frequently incorrect. The Excess Green (ExG) vegetation index and pseudo-thermal representations produced from RGB pictures are two synthetically developed complementary representations that are integrated with RGB imagery in this study’s lightweight multimodal deep learning system to address these issues. Histogram shifting and pseudo-infrared color mapping are used in a reproducible picture alteration pipeline to create the pseudo-thermal modality, which allows for extra visual signals without the need for specific thermal sensors. In order to classify plant diseases while preserving computational efficiency, the suggested framework uses MobileNetV3-Small backbones to extract modality-specific characteristics. This is followed by feature-level fusion. The publicly accessible Ginger Leaf Dataset, which includes RGB pictures of ginger leaves in four different conditions—Damage-Pest, Dehydrated, Healthy, and Leaf-blight—was used for the experiments. For training, validation, and testing, the dataset was split using a stratified 70:15:15 split. Python-based preprocessing procedures were used to create the extra modalities (ExG and pseudo-thermal representations) from the original RGB images. The experimental results show that the combination of the representations with RGB images can enhance the classification performance compared with the unimodal RGB-based models. Ablation experiments are also conducted to examine the contributions of different modalities to the overall categorization accuracy. The experimental results show that plant disease recognition can be improved with the help of efficient computing by combining lightweight convolutional neural networks with computationally generated visual representations.

Why it matches plant phenotyping methodsRGB画像からExG・疑似熱画像を生成し、植物葉の病害状態を分類するマルチモーダル手法が研究の中心であり、アブレーション評価も実施している。

titleHybrid deep learning-based multimodal framework for plant leaf disease classification using RGB, Excess Green (ExG), and pseudo-thermal representations with MobileNetV2
Reproduction assets foundThe paper's phenotyping experiments use the publicly available Ginger Leaf Dataset (RGB leaf images of four ginger leaf conditions), with a public GitHub repository and dataset website. The authors' derived ExG/pseudo-thermal representations and preprocessing scripts are only available upon request, so they do not yet
Dataset · publicor multispectral images IEEE Geosci. Remote Sens. Lett. 2025 10.1109/LGRS.2025.XXXXXXX Ulku, I., Tanriover, O. O. & Akagündüz, E. Cross-band correlation-aware interactive fusion for multispectral images. IEEE Geosci. Remote Sens. Lett. 10.1109/LGRS.2025.XXXXXXX (2025). 10. Wong, J. Ginger Leaf Dataset. GitHub Repository (2023). https://github.com/wongjay1941/Ginger-Leaf-Dataset 11. Bhakta I A novel plant disease prediction model based on thermal images using modified deep convolutional neural network Precis. Agric. 2023 24 23 39 10.1007/s11119-022-09927-x Bhakta, I. et al. A novel plant disease prediction model based on thermal images using modified deep convolutional neural network. Precis.Open asset ↗https://github.com/wongjay1941/Ginger-Leaf-Datasetlines:580-681
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 May 2026Scientific reportsCited by 0 · OpenAlex ↗

Enhancing crop disease recognition framework via vision-language model with cross-attention and gated fusion.

SoybeanMultimodalLeafClassificationStress / disease detectionDisease symptoms / severity

Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate detection of such diseases is crucial for improving both crop yield and quality. While numerous deep learning approaches rely solely on image data for disease identification, they often overlook the complementary value of textual information in enhancing visual analysis. To address this limitation and effectively fuse features from different modalities, we propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition. Our approach utilizes the Zhipu.ai multi-modal model to generate comprehensive textual descriptions of diseased crop leaves, including global description, local lesion description, and color-texture description. These textual descriptions are then encoded into feature embeddings, while visual features are extracted using the ShuffleNet-v2 model as the image encoder. Subsequently, a cross-attention module aligns and fuses the two modalities, and a gated fusion module enables dynamic feature selection during the fusion process. Extensive evaluations on the Soybean Disease and PlantVillage datasets demonstrate that our method outperforms existing image-based models in terms of accuracy. Specifically, our model achieves recognition accuracies of 99.04% and 99.12% on the respective datasets, surpassing the ShuffleNet-V2 model by 1.09% and 2.53%, respectively. These results highlight the effectiveness of Cross-Model learning in integrating visual and textual cues for accurate and efficient disease recognition, offering a scalable solution for crop disease diagnosis.

Why it matches plant phenotyping methods植物葉の病徴を画像と言語情報から認識する融合フレームワークを開発し、複数データセットで既存手法と比較評価しているため、植物フェノタイピング手法が中心である。

abstractwe propose a Cross-Model fusion framework based on a vision-language model that integrates cross-attention and gated fusion mechanisms for crop disease recognition.
Reproduction assets foundThe paper's crop disease recognition experiments use two openly available image datasets, both with explicit public availability statements in the Data Availability section: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). No author analysis code, trained models, or generated text-annotait
Dataset · publicThe datasets utilized in this study are openly accessible. The soybean dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3.Open asset ↗Dryad · 10.5061/dryad.41ns1rnj3html-lines:403-424
Dataset · publicThe plantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-dataset.Open asset ↗Kaggle · plantvillage-datasethtml-lines:403-424
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published16 May 2026Scientific ReportsCited by 0 · OpenAlex ↗

Adaptive fuzzy deep learning with multimodal sensor fusion for enhanced plant disease detection

MultimodalRGB / grayscaleMultispectral / hyperspectralClassificationObject detectionStress / disease detectionDisease symptoms / severity

Abstract Timely and accurate plant disease detection is important for enhancing agricultural productivity and promoting sustainability. The study introduces Multimodal Adaptive Fuzzy-based Deep Neural Network (MAF-DNN) for classification of plant diseases. The proposed method combines fuzzy logic with multimodal data fusion to effectively address the complex interactions and uncertainties in agricultural datasets. The MAF-DNN employs a robust adaptive fuzzy framework with dynamic rule optimization and integrates Hyperspectral Imaging Data (HID) with RGB imaging data to acquire detailed spectral information and high-resolution visual cues for disease classification. The multimodal fusion enhances the model’s ability to capture intricate patterns that relate to plant health, improving the accuracy of disease classification. The experimental results showed that the MAF-DNN outperforms traditional models by achieving an accuracy of 97.8%, precision of 96.5%, recall of 98.2%, and F1-score of 97.3%. Additionally, the adaptive design reduces computational overhead, increases efficiency, and improves scalability for large-scale agricultural applications. The MAF-DNN represents a significant advancement in plant disease classification and provides a robust and efficient solution for precision agriculture.

Why it matches plant phenotyping methods植物病徴を画像から分類するマルチモーダル画像・深層学習手法の開発と性能評価が中心であり、植物の病害状態を直接推定するため。

abstractThe study introduces Multimodal Adaptive Fuzzy-based Deep Neural Network (MAF-DNN) for classification of plant diseases.
Reproduction assets foundThe paper uses two public Kaggle plant disease image datasets (New Plant Diseases Dataset and CCMT Plant Disease Dataset) as its phenotyping inputs and states that the authors' custom MAF-DNN code is publicly available on GitHub, with all three URLs given in the article and matching allowed URLs.
Code · publicThe custom code used to develop and evaluate the proposed Multimodal Adaptive Fuzzy Deep Neural Network (MAF-DNN) framework is publicly available at: https://github.com/skbsangeetha/MAF-DNN-Plant-disease-classificationOpen asset ↗skbsangeetha/MAF-DNN-Plant-disease-classificationhtml-lines:102-118
Code / dataset availability confirmedCrossref · checked 5 Sept 2026
Published30 Apr 2026Plant Science TodayCited by 1 · OpenAlex ↗

AI-driven multi-agent framework for smart irrigation and crop health monitoring in Indian rice and sugarcane farming

RiceSugarcaneAerial / UAVField / plotMultimodalMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldClassificationStress / disease detection

Disease prevention and water management are important to all the crops, particularly rice and sugarcane production in India. The article proposes a reinforcement learning (RL) based intelligent irrigation management system that is capable of optimising water consumption and crop nutrition in response to the changing agricultural climatic conditions. Decentralised reinforcement learning (RL) is used in a network of irrigation agents that utilise soil and microclimate sensor networks to set the terms of water allocation, water use efficiency (WUE) and crop health. At the same time, deep convolutional networks can be used to differentiate between plant stress/disease and leaf images and take applicable proactive actions. It is a framework that incorporates satellite-derived indices (NDVI, EVI, land surface temperature) with local sensor measurements and image-based health measurements through multimodal deep learning. Far-reaching simulations (including Indian climate and crop calendars) demonstrate that the multi-agent system lowers water consumption and preserves the yields and properly notifies stressed plants. The scores of disease detection with plantvillage-based fine-tuned on rice (120 (3 disease types) and 3829 (5 disease types) and sugarcane (2569 images for all disease types, Convolutional Neural Network (CNN) yield results of >98 % accuracy. Crop mapping (rice/sugarcane) Satellite/LSTM-based crop mapping (with Sentinel-1 / Sentinel-2) achieves more than 97 % accuracy. The suggested structure provides a data-driven, scalable system for precision agriculture to enhance the management of irrigation periods and crop health. Simulation experiments show that the RL-based controller can reduce water consumption while preserving optimal soil moisture levels when compared to rule-based irrigation strategies.

Why it matches plant phenotyping methods画像・衛星・センサーを統合して植物ストレス/病害状態を推定するマルチモーダル基盤が提案され、病害検出性能も評価されているため、植物表現型推定が実質的な構成要素である。

abstractdeep convolutional networks can be used to differentiate between plant stress/disease and leaf images
Reproduction assets foundThe paper reports simulation-based experiments using public leaf-image datasets. The only paper-specific public asset explicitly identified is the Kaggle rice leaf diseases dataset (vbookshelf/rice-leaf-diseases) cited as a data source for the rice disease fine-tuning set. No authors' code, trained models, or data dép
Dataset · publicConflict of interest: Authors do not have any conflict of interest 2026 Mar 31). Available from: https://www.kaggle.com/datasets/Open asset ↗Kagglepdf-page:16 lines:1-58
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published21 Apr 2026Scientific reportsCited by 0 · OpenAlex ↗

mIT-CMCA: a cross-modal category alignment framework for robust maize disease identification.

MaizeMultimodalLeafClassificationDisease symptoms / severity

Accurate identification of maize diseases is crucial for safeguarding global food security. Traditional image-based methods often struggle with lighting variations, occlusions, and noise, limiting their robustness and generalisation. Multimodal approaches that integrate visual and textual information have shown promise. However, these methods frequently require manually curated textual descriptions for each image, increasing data collection costs and limiting scalability and practical implementation. To address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA). This approach enforces category-level alignment between image and text modalities, enabling more accurate and interpretable cross-modal mapping. First, we construct cross-modal representations by aligning image and text modalities at the category level within a shared embedding space. Second, inspired by contrastive learning, we introduce a Cross-Modal Category Alignment (CMCA) loss based on category-level textual descriptions, reducing annotation complexity. Finally, we present an Efficient Channel-Spatial Hybrid Attention (CSHA) module that preserves inter-class boundaries while incurring minimal computational overhead, thereby enhancing feature discriminability under complex conditions. Experimental results on the maize subset of the PlantVillage dataset (MPVD) show that mIT-CMCA achieves 99.48% accuracy, 99.28% precision, 99.54% recall, and 99.41% F1-score. These results represent improvements of 0.24%, 0.13%, 0.17%, and 0.15% over the strongest vision-only baseline, MaxViT_tiny. On the self-built Maize Leaf-Field dataset (MLFD), the model achieves 93.67% accuracy, 93.76% precision, 93.67% recall, and 93.71% F1-score. It uses only 8.27 million parameters, which is 71.6% fewer than MaxViT_tiny. Its model size is 32.13 MB, which is 72.3% smaller. The proposed method also outperforms comparative models in robustness experiments under artificially added perturbations. These results demonstrate that mIT-CMCA achieves a favorable balance between accuracy and efficiency, making it suitable for practical agricultural deployment.

Why it matches plant phenotyping methodsトウモロコシ葉画像から病害状態を推定する画像・マルチモーダル手法の開発と性能評価が研究の中心であり、植物病害フェノタイピングに該当する。

abstractwe proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA).
Reproduction assets foundThe paper's maize disease identification analysis code and trained models are explicitly stated as publicly available in a GitHub repository. The phenotype image datasets (MPVD subset and self-built MLFD) are not publicly available and require contacting the corresponding author.
Code · publicCode availability The code and models are available in the GitHub repository at https://github.com/TANGFEILONG626/mIT-CMCA..Open asset ↗TANGFEILONG626/mIT-CMCAhtml-lines:673-695
Code / dataset availability confirmedOpenAlex · checked 5 Sept 2026
Published18 Apr 2026DronesCited by 0 · OpenAlex ↗

drone2report: A Configuration-Driven Multi-Sensor Batch-Processing Engine for UAV-Based Plot Analysis in Precision Agriculture

Aerial / UAVField / plotMultimodalPhotogrammetry / SfM / MVSMultispectral / hyperspectralThermalWhole plant / canopy / plot / fieldClassificationPhysiological trait estimationCalibration / preprocessing

Unmanned aerial vehicles (UAVs) have become indispensable tools in precision agriculture and plant phenotyping, enabling the rapid, non-destructive assessment of crop traits across space and time. Equipped with RGB, multispectral, thermal, and other sensors, UAVs provide detailed information on canopy structure, physiology, and stress responses that can guide management decisions and accelerate breeding programs. Despite these advances, the downstream processing of UAV imagery remains technically demanding. Converting orthomosaics into standardized, biologically meaningful data often requires a combination of photogrammetry, geospatial analysis, and custom scripting, which can limit reproducibility and accessibility across research groups. We present drone2report, an open-source python-based software that processes orthomosaics from UAV flights to generate vegetation indices, summary statistics, derived subimages, and text (html) reports, supporting both research and applied crop breeding needs. Alongside the basic structure and functioning of drone2report, we also present five case studies that illustrate practical applications common in UAV-/drone-phenotyping of plants: (i) thresholding to remove background noise and highlight regions of interest; (ii) monitoring plant phenotypes over time; (iii) extracting information on plant height to detect events like lodging or the falling over of spikes; (iv) integrating multiple sensors (cameras) to construct and optimize new synthetic indices; (v) integrate a trained deep learning network to implement a classification task. These examples demonstrate the tool’s ability to automate analysis, integrate heterogeneous data and models, and support reproducible computation of agronomically relevant traits. drone2report streamlines orthorectified UAV-image processing for precision agriculture by linking orthomosaics to standardized, plot-level outputs. Its modular, configuration-driven design allows transparent workflows, easy customization, and integration of multiple sensors within a unified analytical framework. By facilitating reproducible, multi-modal image analysis, drone2report lowers technical barriers to UAV-based phenotyping and opens the way to robust, data-driven crop monitoring and breeding applications.

Why it matches plant phenotyping methods植物表現型取得のためのUAV画像処理ソフトウェアを開発し、植物高・倒伏などの形質抽出、マルチセンサー統合、再現可能な解析ワークフローを中心的に提示している。

abstractWe present drone2report, an open-source python-based software that processes orthomosaics from UAV flights to generate vegetation indices, summary statistics, derived subimages, and text (html) reports
Reproduction assets foundThe paper explicitly states that the code and data to reproduce its five case studies (thresholding, temporal vegetation indices, height analysis, multi-sensor index optimization, deep learning classification) are publicly available in the authors' GitHub repository, and the DRONE2REPORT software itself is released as
Code · publicThe code and data to reproduce these case studies can be found at https://github.com/ne1s0n/paper-drone2report (accessed on 13 April 2026).Open asset ↗ne1s0n/paper-drone2reportpdf-page:6 lines:1-59
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published27 Mar 2026Plants (Basel, Switzerland)Cited by 1 · OpenAlex ↗

TB-DLossNet: Fine-Grained Segmentation of Tea Leaf Diseases Based on Semantic-Visual Fusion.

Field / plotMultimodalLeafSegmentationDisease symptoms / severity

Camellia oleifera is an economically vital woody oil crop. Its productivity and oil quality are severely compromised by various diseases. Implementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection. Despite recent progress, existing segmentation methods struggle with three primary challenges: semantic ambiguity arising from evolving pathological stages, blurred boundaries due to overlapping lesions, and the high omission rate of micro-lesions. To address these issues, this paper presents TB-DLossNet (Text-Conditioned Boundary-Aware Network with Dynamic Loss Reweighting), a novel segmentation framework based on semantic-visual multi-modal fusion. Leveraging VMamba as the visual backbone, the proposed model innovatively integrates BERT-encoded structured text as an auxiliary modality to resolve visual ambiguities through cross-modal semantic guidance. Furthermore, a boundary enhancement branch is incorporated alongside a multi-scale deep supervision strategy to mitigate boundary displacement and ensure the topological continuity of lesion structures. To tackle the detection of small-scale targets, we designed a dynamic weight loss function conditioned on lesion area, significantly bolstering the model's sensitivity to minute pathological features. Additionally, to alleviate the scarcity of high-quality data, we curated a comprehensive multi-modal dataset encompassing seven typical diseases of Camellia oleifera . Experimental results demonstrate that TB-DLossNet achieves a Mean Intersection over Union (mIoU) of 87.02%, outperforming the state-of-the-art unimodal VMamba and multimodal Lvit by 4.9% and 2.59%, respectively. Qualitative evaluations confirm that our model exhibits lower false-negative rates and superior boundary-fitting precision in heterogeneous field scenarios. Finally, generalization tests on an apple disease dataset further validate the robustness and transferability of the proposed framework.

Why it matches plant phenotyping methods植物病害の病斑を画素レベルで抽出する新規セグメンテーション手法を開発し、データセット整備と性能比較・汎化検証も行っているため、病害状態の画像ベース表現型計測が中心である。

abstractImplementing pixel-level lesion segmentation within complex field environments is crucial for advancing precision plant protection.
Reproduction assets foundThe authors state their code and experimental dataset (the multimodal Camellia oleifera disease segmentation dataset) are publicly available on GitHub, matching an allowed URL.
Code · publicOur code and experimental dataset are available at https://github.com/zzzsq239/TB-1.Open asset ↗zzzsq239/TB-1html-lines:820-841
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published13 Mar 2026Cited by 0 · OpenAlex ↗

A Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis

Field / plotMultimodalWhole plant / canopy / plot / fieldClassificationGrowth / development / phenology

Abstract Phenological monitoring of Actinidia chinensis is critical for optimising operational costs and yield prediction. However, current manual assessment methods are time-consuming, making them impractical for large-scale precision agriculture applications. Most existing phenological datasets focus exclusively on image data without spatial validation. The Multi-Modal Actinidia chinensis Phenology Dataset is composed of (i) 1 665 annotated images of phenological stages from bud to fruit set and (ii) georeferenced videos with systematic manual ground truth of spatial stage distributions. The dataset employs an adapted 17-class BBCH system that consolidates visually similar stages, excludes problematic categories, and introduces generic structural classes to address practical annotation difficulties. Additionally, the data is organised hierarchically across various plant structures, genders, and phenological stages. The annotated images offer versatility for a range of applications, including training data for computer vision models to detect phenological stages. Furthermore, the georeferenced videos facilitate the validation of automated counting algorithms. This combined approach enables plant-level detection accuracy and provides an illustrative methodology for spatial validation that users can extend to additional orchards, promoting the development and benchmarking of automated phenological monitoring systems for precision agriculture applications in kiwifruit production.

Why it matches plant phenotyping methodsキウイフルーツの生育段階を対象とした注釈画像・地理参照動画のデータセットで、植物フェノロジー自動検出の訓練、検証、ベンチマークを目的とする方法論的成果である。

titleA Multi-Modal Dataset for Automated Phenological Stage Mapping in Actinidia chinensis
Reproduction assets foundThe paper is a Data Note describing the Multi-Modal Actinidia chinensis Phenology Dataset, which is explicitly stated to be publicly available on Zenodo with a DOI matching an allowed URL. The dataset contains the paper's own phenotyping assets: 1,665 annotated images with bounding-box phenological labels, georeferened
Dataset · publicThe Multi-Modal Actinidia chinensis Phenology Dataset described in this Data Descriptor is publicly available at Zenodo: https://doi.org/10.5281/zenodo.17371025. This dataset comprises two components: (1) 1 665 JPEG images (1 024 × 1 024 pixels) with corresponding Pascal VOC XML annotation files containing bounding box coordinates and phenological class labels, and (2) 24 MP4 video files (3 840 × 2 160 pixels) with corresponding GPX coordinate files and Excel validation files containing manual ground truth counts.Open asset ↗Zenodo · 10.5281/zenodo.17371025pdf-page:13 lines:1-62
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published2 Mar 2026Plant phenomics (Washington, D.C.)Cited by 1 · OpenAlex ↗

Synthetic-augmented multimodal deep learning fuses dual-angle RGB images and phenology to unlock genotype-informative canopy structural trait in wheat.

WheatField / plotMultimodalRGB / grayscaleWhole plant / canopy / plot / fieldMorphology / geometry measurementGrowth / time-series analysisArchitecture / morphology / geometryGrowth / development / phenologyYield / yield components

The wheat canopy genome harbors abundant yet untapped genetic variation that could be harnessed to enhance yield potential. The green area index (GAI) is a structural metric that reflects the photosynthetically active canopy surface and is closely linked to final grain yield. Current image-based GAI retrieval methods often suffer from signal saturation and coarse structural depiction, constraining downstream genetic analyses. To address this limitation, we constructed a comprehensive image dataset spanning eight field experiments across China and France, encompassing approximately 600 genotypes under six distinct management regimes. Leveraging this diverse data, we developed a multimodal deep-learning framework augmented by simulated-to-realistic (sim2real) synthetic data transfer. This framework fuses nadir and oblique RGB images with accumulated thermal time to produce high-precision, time-series GAI estimates. Validated on independent testing datasets from both China and France, the multimodal approach demonstrated robust performance with an accuracy of R 2 = 0.88 and an RMSE of 0.49 m 2 m -2 , representing an improvement of about 22% over the traditional gap fraction method. In three site-year field experiments involving 565 genotypes, the GAI dynamics derived from the multimodal approach showed higher broad-sense heritability (0.20-0.48) than those from the gap fraction approach (0.02-0.13) and stronger genotypic correlations with yield (0.19-0.40 versus 0.09-0.31). Furthermore, genetic analysis confirmed the biological fidelity of the estimated traits, identifying loci that co-localize with known architectural regulators such as Rht-D1 , TaTB1-4D , and TaBGC1-4D . Consistently, the multimodal-derived phenotypes were specifically enriched in cell-wall remodeling and hormonal signaling pathways (e.g., brassinosteroid) that directly regulate canopy expansion. Overall, the proposed method offers a powerful tool for unlocking genetic gain in canopy architecture and accelerating canopy-targeted wheat improvement.

Why it matches plant phenotyping methodsデュアルアングルRGB画像と熱時間を統合してGAIを推定する深層学習法を開発し、独立データで検証しているため、植物形質取得法が研究の中心です。

abstractwe constructed a comprehensive image dataset spanning eight field experiments across China and France
Reproduction assets foundThe paper publicly releases its pre-trained multimodal GAI-estimation model weights and inference code on Hugging Face, directly reproducing this paper's phenotyping analysis. The raw image and phenology datasets are not public and require contacting the authors.
Code · publicThe pre-trained model weights, inference code, and usage instructions are publicly available in the Hugging Face repository at https://huggingface.co/PheniX-Lab/GAI-Estimation/tree/main .Open asset ↗PheniX-Lab/GAI-Estimationlines:259-277
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published2 Mar 2026Data in briefCited by 0 · OpenAlex ↗

Seasonal collection of in situ optical and thermal images dataset and meteorological measurements over an Indian semi-arid rice crop.

RiceField / plotMultimodalMultispectral / hyperspectralThermalWhole plant / canopy / plot / fieldCalibration / preprocessingLeaf traitsPlant / canopy heightPlant / canopy temperature

This article describes a multi-sensor dataset collected during the TIRAMISU (Thermal InfraRed Anisotropy Measurements in India and Southern eUrope) campaign at the Nawagam research site in Gujarat, India, during the 2023 monsoon season. The objective was to acquire continuous ground-based optical and thermal measurements over a homogeneous rice canopy across different crop growth stages. The dataset integrates several complementary components. Thermal data were acquired with an Optris longwave infrared camera (8-14 µm) at high temporal resolution, capturing canopy temperature dynamics throughout the diurnal cycle. Optical data were obtained with a Micasense RedEdge-M multispectral sensor, providing imagery in Blue, Green, Red, RedEdge, and Near-Infrared bands with radiometric corrections. An Apogee radiometer supplied reference radiometric temperature. Meteorological measurements included air temperature, humidity, wind speed and direction, and net radiation. Ancillary field measurements comprised Leaf Area Index (LAI), plant height, emissivity sampling, hyperspectral observations, and crop stage information. The datasets are provided with metadata and processing workflows, including calibration procedures for optical reflectance and thermal radiance. Together, these components form a comprehensive record of canopy-atmosphere interactions over a homogeneous rice field. The datasets can support research on optical and thermal directional anisotropy, canopy radiative transfer, emissivity characterization, and crop biophysical parameter estimation. In addition, they are relevant for applications in vegetation monitoring, agricultural water stress assessment, and surface energy balance studies. By combining optical, thermal, and meteorological observations, the resource is suited for multidisciplinary investigations in remote sensing, agronomy, and environmental sciences.

Why it matches plant phenotyping methods光学・熱画像、校正手順、処理ワークフロー、LAIや草丈などの植物形質を含む再利用可能な作物キャノピーデータセットが研究の中心であり、植物表現型取得基盤として適格。

abstractThe dataset integrates several complementary components.
Reproduction assets foundThe paper is a Data in Brief describing the TIRAMISU rice-canopy dataset (thermal/multispectral images, meteorological, ancillary LAI/height, hyperspectral, emissivity) publicly deposited at doi.org/10.6096/1028, including processing scripts (Thermal_CSV_to_Image.py, MicaSense notebook) for reproducibility.
Dataset · publicRepository name: Optical, Thermal Infrared, and Meteorological Dataset from the Thermal InfraRed Anisotropy Measurements in India and Southern eUrope (TIRAMISU) Rice Canopy Experiment Data identification number: doi.org/10.6096/1028 Direct URL to data: https://doi.org/10.6096/1028 Instructions for access: Publicly accessible repository; representative subsets provided with metadata and processing scripts. Related research article Pinnepalli, C., Roujean, J.-L., Irvine, M., et al. [ 1 ]. Measuring and modelling directional effects in the frame of TIRAMISU. ISPRS Annals, X–3–2024 , 325–330. https://doi.orgOpen asset ↗doi.org · 10.6096/1028lines:49-77
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published28 Feb 2026Scientific reportsCited by 2 · OpenAlex ↗

Design and implementation of a deep learning framework for automated crop classification and health diagnosis in precision agriculture.

MaizePotatoWheatAerial / UAVMultimodalStress / disease detectionStress response / tolerance

This paper presents a three-phase deep learning framework comprising (i) multi-modal data acquisition from drones and satellites, (ii) standardized pre-processing including interpolation for missing temporal data, and (iii) CNN-based feature extraction for real-time health classification. This framework relies on a mathematical model based on neural networks that classifies and detects the condition of agriculture, removing the reliance on manual tasks and subjective diagnosis. This paper focuses on three main aspects of our framework: data acquisition, training and prediction. Data is collected using sensors like drones, cameras, and satellite imagery and is pre-processed to filter out noise and improve quality. The training part uses CNN to learn features from the data and become more meaningful. The prediction part of the task classifies, and diagnoses crop health through the trained model using the features. The framework accuracy for crops such as maize, potato, and wheat has been tested and yielded over 90% accuracy. The novelty of this work resides in the development of a multi-modal deep learning architecture that fuses macro-scale satellite imagery with micro-scale drone and IoT sensor data to improve diagnostic reliability. The framework was validated on a multi-source agricultural dataset using a 70% training, 15% validation, and 15% testing protocol. Experimental results demonstrate an accuracy exceeding 90% for staple crops. Using this framework can increase the visibility and quality of information maintained for crop health and improve the decision-making routine of farmers in real time. Additionally, automation of this process can significantly reduce labor costs and increase productivity per crop. Implementing this framework can contribute to precision agriculture and sustainable management practices.

Why it matches plant phenotyping methods作物の健康状態を植物の表現型・状態として推定するマルチモーダル画像・センサ基盤と深層学習手法を開発し、複数作物・データセットで検証しているため、方法が中心である。

abstractThis paper presents a three-phase deep learning framework comprising (i) multi-modal data acquisition from drones and satellites, (ii) standardized pre-processing including interpolation for missing temporal data, and (iii) CNN-based feature extraction for real-time health classification.
Reproduction assets foundThe article's Data availability section points to a public Kaggle dataset used for the crop classification/health diagnosis experiments, matching an allowed URL. No code or model checkpoints are disclosed.
Dataset · publicript. The research work was guided by Dr. B.D.K.P. The Corresponding author Shshank Chaube collaborated for review and supervision. All authors reviewed the manuscript. Funding Open access funding provided by Symbiosis International (Deemed University). No funds, grants, or other support was received. Data availability Dataset: https://www.kaggle.com/datasets/bhagvendersingh/precision-agriculture-dataset . Declarations Competing interests The authors declare no competing interests. Ethical approval This article does not contain any studies with human participants or animals performed by any of the authors. References 1. Mohyuddin, G. et al. Evaluation of machine learning approaches for preciOpen asset ↗kaggle · bhagvendersingh/precision-agriculture-datasetlines:473-545
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 15 Sept 2026
Published6 Feb 2026Plant PhenomicsCited by 1 · OpenAlex ↗

Fine-grained 3D rice phenotyping via multi-scale NeRF and multimodal segmentation.

RiceField / plotMultimodalNeRF / 3D Gaussian SplattingLiDAR / point cloudSeed / grainWhole plant / canopy / plot / fieldMorphology / geometry measurement2D/3D reconstructionSegmentation

Fine-grained 3D phenotypic analysis of rice plays a vital role in rice breeding and yield estimation. However, a comprehensive rice data acquisition and segmentation pipeline is still lacking. While Neural Radiance Fields (NeRF) have shown impressive results in crop-level 3D reconstruction, their high sensitivity to data volume and camera viewpoints often leads to reconstruction failures for rice. In addition, the large-scale rice point clouds, coupled with heavy occlusion and visual similarity among grains, pose significant challenges for fine-grained trait extraction. To address the challenge of reconstructing rice point clouds under low-quality data conditions, we propose a novel method named Multi-Scale NeRF(MSNeRF). This method incorporates a structure-detail collaborative reconstruction mechanism and a dynamic initialization density scheduling strategy. Furthermore, we introduce a multimodal and multitask rice dataset (MMR) as a benchmark resource for future research. For rice point cloud segmentation, we develop Vision Rice Knowledge Graph Network(VRKGNet), which comprises an image segmentation module, a projection module, and a point cloud segmentation module enhanced with a Transformer to enlarge the receptive field. VRKGNet performs standalone point cloud segmentation and integrates image segmentation results from multiple viewpoints as prior knowledge to enhance semantic and instance-level segmentation. Extensive experiments demonstrate that MSNeRF achieves high-fidelity point cloud reconstruction with as few as 10 viewpoints. VRKGNet achieves superior rice plant segmentation with a semantic segmentation mIoU of 88.79% and an instance segmentation AP 25 of 84.55%, outperforming mainstream algorithms.

Why it matches plant phenotyping methods米の3D形質取得・再構成・分割を中核とする手法開発であり、データセット/ベンチマークも提供しているため、植物フェノタイピング手法文献に該当する。

abstractwe propose a novel method named Multi-Scale NeRF(MSNeRF)
Reproduction assets foundThe paper's authors explicitly state that the source code for MSNeRF and VRKGNet is publicly available on GitHub with testing scripts and test cases to reproduce the main results. The MMR dataset itself is only available upon request from the corresponding author, so it does not qualify as a public asset.
Code · publicof Hefei Artificial Intelligence Breeding Accelerator Co. Ltd. ( NB2024005-02 ). Data availability The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper: • MSNeRF : https://github.com/qfwysw/MSNeRF.git • VRKGNet : https://github.com/qfwysw/VRKGNet.git The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author. Declaration of competing interest The authors declare that they have no known competing financial iOpen asset ↗https://github.com/qfwysw/MSNeRF.gitlines:620-663
Code · publicator Co. Ltd. ( NB2024005-02 ). Data availability The source code for the proposed methods, MSNeRF and VRKGNet, is publicly available on GitHub. The released repositories contain testing scripts and test cases used to reproduce the main results presented in this paper: • MSNeRF : https://github.com/qfwysw/MSNeRF.git • VRKGNet : https://github.com/qfwysw/VRKGNet.git The datasets used in the experiments are available from the corresponding author upon reasonable request. For access or further inquiries, please contact the corresponding author. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could haveOpen asset ↗https://github.com/qfwysw/VRKGNet.gitlines:620-663
Code / dataset availability confirmedCrossref · Europe PMC · checked 5 Sept 2026
Published30 Jan 2026PlantsCited by 0 · OpenAlex ↗

Introducing Concurrent Imaging and Unidimensional Analytics for Plant Stress Responses.

MultimodalSegmentationStress / disease detectionStress response / tolerance

Advancements in phenotyping technologies, including object imaging, high-throughput monitoring, and soft computing, are pivotal for understanding plant responses to environmental stresses. These technologies enable detailed analyses of morphological, physiological, and structural adaptations under abiotic and biotic stresses, such as drought. Current work using multimodal and multi-perspective image processing methods can capture the essential processes that enhance plant resilience and counteract stress by identifying morphological and biochemical indicators. However, the dynamic and complex nature of plant responses poses multiple challenges for generating precise analytics and descriptors of evolving phenotypes. This work introduces analytics for concurrent imaging, adopting the underlying principle of cosegmentation to create taxonomies for new phenotypes. Here, unidimensional refers to the concurrent analysis of multiple images within a single phenotyping dimension: temporal, modal, or perspective, rather than combining information across dimensions. The proposed unidimensional phenotypes integrate concurrent images within individual temporal, modal, or perspective dimensions to capture dynamic morphological and physiological responses that are not observable with conventional single-image or cumulative metrics. Within a high-throughput imagery production system, these phenotypes enable more nuanced quantification of phenotypic changes, leveraging the strengths of simultaneous image analysis to enhance insight into plant adaptations. This workflow aligns with the investigation of plants’ adaptive strategies under abiotic stress and provides quantitative indicators of plant health under adverse environmental conditions.

Why it matches plant phenotyping methods植物の同時画像解析とコセグメンテーションに基づく新しい表現型抽出・定量化ワークフローを提案しており、植物フェノタイピング手法が中心である。

abstractThis work introduces analytics for concurrent imaging, adopting the underlying principle of cosegmentation to create taxonomies for new phenotypes.
Reproduction assets foundThe paper's Data Availability Statement explicitly states that the SIMID and SIPID image datasets created and used in this study are publicly available on Zenodo (DOI 10.5281/zenodo.17400167), which is an allowed URL. These are the paper-specific plant phenotyping imagery inputs (buckwheat and sunflower under control/d
Dataset · publicThe SIMID and SIPID dataset utilized and created in this study is publicly available and accessible at the following link: https://doi.org/10.5281/zenodo.17400167Open asset ↗Zenodo · 10.5281/zenodo.17400167lines:174-216
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published29 Dec 2025Scientific reportsCited by 5 · OpenAlex ↗

Reinforcement learning based dynamic vegetation index formulation for rice crop stress detection using satellite and mobile imagery.

RiceField / plotMultimodalRGB / grayscaleMultispectral / hyperspectralWhole plant / canopy / plot / fieldClassificationStress / disease detectionStress response / tolerance

Timely crop stress detection is essential for safeguarding yields and promoting sustainable agriculture. Traditional vegetation indices (e.g., NDVI, EVI) are widely used but remain static, crop-agnostic, and often insensitive to early stress signals. This study proposed RL-VI, a reinforcement learning-based framework that dynamically formulates vegetation indices optimized for rice stress detection. Unlike existing methods, RL-VI integrates Sentinel-2 multispectral imagery with smartphone-captured RGB data, creating the first cross-platform environment where vegetation indices are learned rather than predefined. The reinforcement learning agent adaptively selects stress-sensitive spectral band combinations guided by classification rewards. Experiments on real-world rice fields in Tamil Nadu, India, and benchmark datasets (Indian Pines, wheat salt stress) show that RL-VI achieves an overall accuracy of 89.4% and F1-score of 0.88, outperforming static and machine-learned indices by up to 12%. Importantly, RL-VI enables early stress detection up to 10 14 days before visible symptoms, providing actionable lead time for intervention. The proposed framework is computationally lightweight and scalable to UAV or edge devices, offering a farmer-ready tool for precision agriculture, bridging field-level mobile sensing with satellite monitoring for low-cost, real-time crop health management. Statistical validation using ANOVA (F = 88.24, p < 0.001) and pairwise t-tests (p < 0.001) confirmed RL-VI's superiority, while SHAP analyses emphasized the physiological significance of red-edge and SWIR bands in stress discrimination.

Why it matches plant phenotyping methods植物ストレス状態を推定する動的植生指数と強化学習フレームワークを開発し、実圃場・ベンチマークデータで性能検証しているため、フェノタイピング手法が中心である。

abstractThis study proposed RL-VI, a reinforcement learning-based framework that dynamically formulates vegetation indices optimized for rice stress detection.
Reproduction assets foundThe paper publicly releases its authors' field-captured mobile RGB rice canopy dataset on Kaggle and its full RL-VI analysis code (RL formulation, preprocessing, VI computation, training, evaluation) on GitHub. Sentinel-2 imagery and benchmark datasets are third-party public sources, not paper-specific deposits.
Dataset · publicThe Mobile RGB dataset, consisting of field-captured rice canopy images collected by the authors at Polur, Tamil Nadu, India, is publicly available on Kaggle under a CC BY-NC 4.0 license (DOI: [https://doi.org/10.34740/kaggle/dsv/14105754](https:/doi.org/10.34740/kaggle/dsv/14105754)).Open asset ↗Kaggle · 10.34740/kaggle/dsv/14105754html-lines:616-683
Code · publicAll custom code developed for this work including the RL-VI (Reinforcement Learning–based Vegetation Index) formulation algorithm, image preprocessing scripts, vegetation index computation modules, model training pipelines, and evaluation routines is openly accessible in a public GitHub repository. The code is available without restriction for non-commercial research use and fully available at Github Repository (https://github.com/Poornisrm/Vegetation-Index.git).Open asset ↗GitHub · Poornisrm/Vegetation-Indexhtml-lines:684-711
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published25 Dec 2025Scientific reportsCited by 1 · OpenAlex ↗

Robust Fall Army Worm detection in maize using multimodal RGB and thermal image fusion.

MaizeMultimodalRGB / grayscaleThermalWhole plant / canopy / plot / fieldClassificationDisease symptoms / severity

Effective pest and disease detection plays a crucial role in minimizing crop losses and improving decision-making in precision agriculture. Among the most destructive pests affecting maize crops globally is the Fall Army Worm (FAW), known for its rapid spread and high impact on yield. Existing detection practices often rely on manual scouting, which can be inefficient, labour intensive and prone to human error. This study proposes a novel deep learning based framework for the automatic classification of FAW infested and healthy maize crops by integrating RGB and thermal image modalities. The core objective is to enhance detection accuracy through multimodal image fusion. A hybrid DNN-ViT model is introduced, combining two complimentary pipelines: (i) feature-level fusion, where CNN extracted features from RGB and thermal images are fused and classified using a Deep Neural Network (DNN) and (ii) image-level fusion, where a 6 channel RGB-thermal image is directly processed using a modified Vision Transformer (ViT). Experimental results demonstrate that the fused model achieved superior performance with an accuracy of 0.98, precision, recall and F1-score of 0.98 and AUC-ROC of 0.98 on the test set, outperforming models trained on RGB-only, thermal-only and unfused data. The ablation study confirms the effectiveness of multimodal fusion, with the no-fusion model showing significantly lower performance (accuracy-0.60 and AUC-ROC-0.67). This work highlights the benefits of integrating complementary data sources for robust crop health monitoring. Future research will explore enhanced fusion strategies, environmental robustness and field level deployment to validate the model's practical applicability.

Why it matches plant phenotyping methodsRGB・熱画像融合によるFAW被害・健全状態の画像判定モデルを開発し、融合方式や性能を比較検証しているため、植物の健康状態を取得する方法が中心である。

abstractThis study proposes a novel deep learning based framework for the automatic classification of FAW infested and healthy maize crops by integrating RGB and thermal image modalities.
Reproduction assets foundThe paper's paired RGB/thermal maize FAW image dataset is publicly deposited on Figshare (part of a peer-reviewed data publication), and the authors' custom Python analysis code is released as a public supplementary file (Supplementary Code.zip) with explicit availability language. The Figshare URL matches an allowed,
Dataset · publicThe dataset has been made publicly available in the Figshare Data repository as a part of a peer reviewed data publication54. Detailed information on data acquisition, sensor specifications, environmental conditions and annotation protocols is provided in the associated data article. The dataset can be accessed at: https://figshare.com/s/677d2384ba6e02db9230 (10.6084/m9.figshare.28388018).Open asset ↗Figshare · 10.6084/m9.figshare.28388018html-lines:324-345
Code · publicThe custom python code developed for this study is available as supplementary file (“Supplementary Code.zip”) and includes all scripts necessary to reproduce the multimodal feature fusion, image-level fusion and ablation experiments described in the manuscript. The dataset used is publicly available on Figshare. All dependencies are listed within the code file. Readers can execute the python script to reproduce the reported results.Open asset ↗html-lines:324-345
Code / dataset availability confirmedEurope PMC · checked 5 Sept 2026
Published18 Dec 2025Scientific reportsCited by 3 · OpenAlex ↗

TomatoRipen-MMT: transformer-based RGB and NIR spectral fusion for tomato maturity grading.

TomatoGreenhouseMultimodalRGB / grayscaleMultispectral / hyperspectralFruitClassificationSegmentationGrowth / development / phenology

Computer vision and multispectral imaging have increasingly become essential tools in modern precision agriculture. Accurate ripeness assessment is critical for yield optimization, reducing post-harvest losses, and enabling automated harvesting systems. However, traditional RGB-based approaches struggle to differentiate subtle maturity changes, and existing solutions often fail under varying lighting, occlusion, or cultivar-specific conditions. To address these challenges, this study focuses on the integration of complementary spectral cues for reliable tomato ripeness evaluation. The work utilizes a curated RGB-NIR tomato dataset comprising 224 hyperspectral samples, processed into aligned multimodal image pairs with balanced ripeness categories.The proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues. The novelty of the methodology lies in the dynamic cross-attention mechanism, which learns inter-modal dependencies between RGB and NIR signals for enhanced ripeness interpretation. Performance metrics including accuracy, precision, recall, F1-score, mIoU, and AUC were used to comprehensively evaluate the system. Experimental results demonstrate that TomatoRipen-MMT significantly outperforms all baseline RGB-only, NIR-only, and fusion methods, achieving 94.8% classification accuracy and 82.6% mIoU. These findings establish the effectiveness of multimodal Transformers for robust, high-precision fruit maturity assessment in controlled and greenhouse environments.

Why it matches plant phenotyping methodsトマト果実の成熟度という植物器官の状態を、RGB・NIR画像融合とTransformerで推定する手法を開発・評価しており、フェノタイピング手法が中心です。

abstractThe proposed TomatoRipen-MMT model employs a multimodal Transformer framework with dual encoders, cross-spectral attention, and a joint decoder to fuse spatial and biochemical cues.
Reproduction assets foundThe paper's phenotyping analysis is built on a publicly available USDA/NAL hyperspectral tomato dataset, explicitly linked in the Data Availability statement with an exact URL match. No author code or model checkpoints are disclosed.
Dataset · publicThe dataset analyzed in this study is publicly available at the https://agdatacommons.nal.usda.gov/articles/dataset/Data_from_b_Hyperspectral_Imaging_Analysis_for_Early_Detection_of_Tomato_Bacterial_Leaf_Spot_Disease_b_/26046328.Open asset ↗26046328html-lines:1038-1053
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published3 Dec 2025Scientific dataCited by 2 · OpenAlex ↗

Maps of forest vertical structure for Colombia, a megadiverse country.

MultimodalLiDAR / point cloudMultispectral / hyperspectralWhole plant / canopy / plot / field2D/3D reconstructionArchitecture / morphology / geometryPlant / canopy height

Vegetation vertical structure refers to the 3D distribution of vegetation aboveground biomass. Vegetation vertical structure of tropical forests influences other ecological and environmental variables that are essential for the functioning of the ecosystems. Integrating over 5.9 million Globel Ecosystem Dynamics Investigation (GEDI) LiDAR (Light Detection and Ranging) footprints, multispectral, and synthetic aperture radar (SAR) imagery, we built five national maps at 25 m resolution of five forest structural metrics for Colombia, South America, for the year 2020. We mapped canopy height, the height of half the cumulative returned energy from GEDI (RH50), total canopy cover, foliage height diversity, and total plant area index. The resulting maps tended to have the highest errors in the Amazon and Andean regions. Total cover had the highest relative error. Interrelationship curves between forest structural metrics of GEDI footprints are maintained across mapped metrics, indicating that the predictive models preserve structural relationships observed in GEDI data. Due to the medium-high spatial resolution and national coverage of the forest structural maps presented in this work, these maps will be useful for evaluating and mapping other ecological variables and conservation priorities in Colombia.

Why it matches plant phenotyping methodsGEDI LiDAR・マルチスペクトル・SARを統合し、森林キャノピー高、被覆率、葉群高多様性、植物面積指数などの植物構造形質を全国規模で推定・検証することが中心であり、単なる生態学的応用ではない。

abstractIntegrating over 5.9 million Globel Ecosystem Dynamics Investigation (GEDI) LiDAR (Light Detection and Ranging) footprints, multispectral, and synthetic aperture radar (SAR) imagery, we built five national maps at 25 m resolution of five forest structural metrics for Colombia, South America, for the year 2020.
Reproduction assets foundThe paper's resulting forest vertical structure maps (CH, COVER, FHD, PAI, RH50 for Colombia, 2020) are publicly available on Zenodo and via Google Earth Engine assets, and the authors' analysis code is publicly available on GitHub. These are paper-specific, public, actionable assets.
Code · publicCode availability The code is publicly accessible on Github76: https://github.com/CamiloFaguaUNAL/Forest_Structure_Colombia.Open asset ↗GitHubhtml-lines:731-755
Code / dataset availability confirmedCrossref · OpenAlex · Europe PMC · checked 14 Sept 2026
Published1 Dec 2025Plant PhenomicsCited by 17 · OpenAlex ↗

Deep learning for three-dimensional (3D) plant phenomics

MultimodalLiDAR / point cloudAnnotation / quality controlClassificationObject detectionCalibration / preprocessingSegmentationTracking

Plant phenomics, the comprehensive study of plant phenotypes, has gained prominence as a vital tool for understanding the intricate relationships between genotypes and the environment. Image-based plant phenomics has progressed rapidly, and three-dimensional (3D) phenotyping is a valuable extension of traditional 2D phenomics. However, the increased data dimensionality poses challenges to feature extraction and phenotyping. In recent decades, deep learning has led to remarkable progress in revolutionizing 3D phenotyping. Therefore, this review highlights the importance of using deep learning in 3D plant phenomics. It systematically overviews the capabilities of deep learning for 3D computer vision, covering 3D representation, classification, detection and tracking, semantic segmentation, instance segmentation, and generation. Additionally, deep learning techniques for 3D point preprocessing (e.g., annotation, downsampling, and dataset organization) and various plant phenotyping tasks are discussed. Finally, the challenges and perspectives associated with deep learning in 3D plant phenomics are summarized, including (1) benchmark dataset construction by using synthetic datasets and methods such as generative artificial intelligence and unsupervised or weakly supervised learning; (2) accurate and efficient 3D point cloud analysis by leveraging multitask learning, lightweight models, and self-supervised learning; and (3) deep learning for 3D plant phenomics by exploring interpretability, extensibility, and multimodal data utilization. The exploration of deep learning in 3D plant phenomics is poised to spur breakthroughs in a new dimension of plant science.

Why it matches plant phenotyping methods3D植物フェノミクスにおける深層学習手法を体系的にレビューしており、植物形質の抽出・推定手法が中心である。

abstractTherefore, this review highlights the importance of using deep learning in 3D plant phenomics.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Dataset · publicThe dataset can be downloaded from https://github.com/Jinlab-AiPhenomics/Mazie3D.Open asset ↗Jinlab-AiPhenomics/Mazie3Dhtml-lines:332-336
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published22 Nov 2025Data in briefCited by 0 · OpenAlex ↗

Phenology and health of Stenocereus Queretaroensis : A multimodal dataset combining multispectral imagery and spectrophotometry.

Field / plotMultimodalMultispectral / hyperspectralRaman / spectroscopyWhole plant / canopy / plot / fieldCalibration / preprocessingGrowth / development / phenology

This data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo), a native species from the arid and semi-arid regions of Southern Zacatecas, Mexico. In particular, Stenocereus spp. are important cacti in the region due to its nutritional properties, role as an economic resource, and cultural significance.It is worth emphasising that these cacti traditionally grow wild (i.e., without deliberate cultivation); accordingly, controlled cultivation is uncommon and remains understudied. With the aim of producing a formal, comprehensive analysis and compendium, the data were collected across multiple phenological stages to provide a complete representation of the plant development cycle, from vegetative growth through to fruiting. To achieve this, the collection process combined high-resolution multispectral imaging with field spectrometry in the 400-700 nm range. Standardized acquisition protocols were applied in field conditions to capture consistent reflectance data, and environmental variables such as illumination, temperature, and geographic coordinates were recorded for each session to ensure reproducibility. The dataset integrates several components: (i) multispectral images that provide spatial information on canopy and structural characteristics, (ii) field spectral signatures with detailed reflectance values for each sampled plant, and (iii) metadata describing phenological stage, acquisition date and time, environmental conditions, and equipment settings. For subsequent analysis, data was preprocessed and normalized to enable reliable comparisons between growth stages and across acquisition sessions, resulting in a clean, structured resource ready for computational analysis. In this regard, this dataset has been organized to facilitate its direct application across multiple research and development contexts. Specifically, potential applications include the training and validation of machine learning and computer vision models for automated phenological stage classification, harvest time estimation, and development of species-specific vegetation indices. Moreover, owing to its standardized design, the resource can serve as a benchmark for comparing methods, validating algorithms, and supporting reproducible workflows in precision agriculture and remote sensing. Beyond Stenocereus queretaroensis, the documented acquisition and preprocessing methodology can be replicated or adapted to generate similar multimodal datasets for other climate-resilient crops, particularly those cultivated in arid and semi-arid regions. This could enable comparative analyses across species and provide a reference for extending multimodal sensing approaches to underrepresented plants of ecological and economic importance.

Why it matches plant phenotyping methods植物の生育段階・生理・構造特性を対象に、標準化されたマルチスペクトル画像とフィールド分光データを収集・前処理した再利用可能なデータセットであり、ベンチマークやアルゴリズム検証を目的とするため、フェノタイピング手法が中心です。

abstractThis data article presents a multimodal, non-invasive dataset documenting the physiology and growth stages of Stenocereus queretaroensis (pitayo)
Reproduction assets foundThe paper's own multimodal phenotyping dataset (multispectral/RGB images, spectral signatures, NDVI products, metadata, and example MATLAB scripts) is publicly deposited on Mendeley Data with explicit direct URL and DOI.
Dataset · public) at ∼1750 m a.s.l., under semi-arid temperate conditions with spring temperatures ranging 20–33°C. The data were collected from the Unit Academic of Electrical Engineering Plantel Jalpa. Data accessibility Repository name: Multimodal_Cactaceae_Dataset_25 Data identification number: doi:10.17632/skw8tjc82f.1 Direct URL to data: https://data.mendeley.com/datasets/skw8tjc82f/1 Instructions for accessing these data: click on the direct URL to obtain the multimodal data from Mendeley Dataset Repository. Related research article None 1. Value of the Data • These data provide a unique, non-invasive resource for studying Stenocereus spp. physiology. The integrated collection of high-resolution multOpen asset ↗Mendeley Data · doi:10.17632/skw8tjc82f.1lines:32-58
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 6 Sept 2026
Published19 Nov 2025Frontiers in plant scienceCited by 7 · OpenAlex ↗

GAE-YOLO: a lightweight multimodal detection framework for tomato smart agriculture with edge computing

TomatoMultimodalStereoFruitObject detectionVisualization / data managementGrowth / development / phenologyFruit / seed / panicle traitsYield / yield components

Introduction The advancement of smart agriculture has witnessed increasing applications of computer vision in crop monitoring and management. However, existing approaches remain challenged by high computational complexity, limited real-time capability, and poor multi-task coordination in tomato cultivation scenarios. Methods To address these limitations, an intelligent tomato management system is proposed based on the Ghost-based Adaptive Efficient You Only Look Once (GAE-YOLO) algorithm. The lightweight architecture of the GAE-YOLO framework is achieved through the replacement of standard convolutional layers with Ghost Convolution (GhostConv) modules, while detection accuracy is significantly improved by the integration of both AReLU activation functions and Effective Intersection over Union (E-IoU) loss optimization. The system, implemented on a Jetson TX2 embedded platform, also incorporates ZED stereo vision for 3D localization and a PyQt6-based visualization platform. Results When implemented on Jetson TX2, the system achieving 93.5% mean Average Precision at 50% intersection over union (mAP@50) at 10.2 frames per second (FPS), which can be optimized to 27 FPS by employing TensorRT acceleration and 720p resolution for scenarios demanding higher throughput. Furthermore, it establishes standardized assessment systems for tomato maturity and yield prediction, and offers integrated modules for disease diagnosis and agricultural large language model consultation. Discussion This work establishes a new paradigm for edge computing in agriculture while providing critical technical support for smart farming development.

Why it matches plant phenotyping methodsトマトの成熟度・収量予測および病害診断を含む画像・3Dビジョン基盤を開発し、エッジ環境で性能評価しているため、植物表現型取得が中心的な研究である。

abstractan intelligent tomato management system is proposed based on the Ghost-based Adaptive Efficient You Only Look Once (GAE-YOLO) algorithm
Reproduction assets foundThe paper's data availability statement explicitly states that the data and code supporting the study are publicly available on GitHub at the authors' repository (GAE-YOLO), which matches an allowed URL. This qualifies as a paper-specific public code asset for the tomato detection/phenotyping analysis.
Code · publicThe data and code supporting this study are publicly available at GitHub under the following links: https://github.com/NSSCk/GAE-YOLO .Open asset ↗NSSCk/GAE-YOLOlines:756-834
Code / dataset availability confirmedEurope PMC · checked 6 Sept 2026
Published5 Nov 2025Sensors (Basel, Switzerland)Cited by 3 · OpenAlex ↗

Early Detection of Jujube Shrinkage Disease by Multi-Source Data on Multi-Task Deep Network.

MultimodalRGB / grayscaleMultispectral / hyperspectralFruitClassificationStress / disease detectionDisease symptoms / severity

In the arid cultivation region of Xinjiang, China, shrinkage disease severely compromises the quality, yield, and market value of jujube. Published research has achieved high accuracy in detecting larger lesions using RGB imaging and hyperspectral imaging (HSI). However, these methods lack sensitivity in detecting early and subtle symptoms of disease. In this study, a multi-source data fusion strategy combining RGB imaging and HSI was proposed for non-destructive and high-precision detection of early-stage jujube shrinkage disease. Firstly, a total of 317 fruits of the 'Junzao' cultivar were collected during multiple stages of natural infection, covering early-stage shrinkage disease detection across different growth stages, including both green and mature red fruits. Secondly, morphological features were extracted from RGB images in multiple dimensions, while a three-stage feature selection strategy combining Principal Component Analysis (PCA), the Successive Projections Algorithm (SPA), and the Genetic Algorithm (GA) was implemented to identify four key wavelengths from HSI. Thirdly, a hybrid convolutional neural network-multilayer perceptron (CNN-MLP) architecture was constructed, with dynamic feature weighting employed to achieve effective multimodal fusion and optimize detection performance. Experimental results demonstrated that compared to the MLP and CNN models, the proposed method achieved approximately 8.0% and 5.4% improvements in accuracy and 38.6% and 32.4% improvements in F1 scores, respectively. It offers a robust and scalable solution for early disease detection and postharvest quality assessment in jujube production.

Why it matches plant phenotyping methodsRGB画像・HSIから果実の病斑形態と分光特徴を抽出し、マルチモーダル深層学習で植物病害状態を検出する手法の開発・性能評価が中心であるため。

abstracta multi-source data fusion strategy combining RGB imaging and HSI was proposed for non-destructive and high-precision detection of early-stage jujube shrinkage disease.
Reproduction assets foundThe paper's Data Availability Statement points to a public GitHub repository containing the study's dataset (RGB images and hyperspectral data of jujube fruits). No separate analysis code availability is stated, but the deposited dataset is a paper-specific, publicly actionable asset.
Dataset · publicThe data from this study are publicly available. The dataset is available at https://github.com/2484733079/Early-detection-of-Jujube-Shrinkage-Disease-by-Multi-source-Data-on-Multi-task-Deep-Network.git (accessed on 13 October 2025).Open asset ↗2484733079/Early-detection-of-Jujube-Shrinkage-Disease-by-Multi-source-Data-on-Multi-task-Deep-Networklines:365-367
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 14 Sept 2026
Published21 Oct 2025Plant PhenomicsCited by 10 · OpenAlex ↗

PlantIF: Multimodal semantic interactive fusion via graph learning for plant disease diagnosis.

MultimodalClassificationStress / disease detectionDisease symptoms / severity

Plant diseases remain a major constraint on crop productivity, requiring timely and accurate diagnostic approaches to secure agricultural yields. While existing automated diagnosis methods primarily rely on image data and achieve notable results, their performance often declines in complex field environments with noise and interference. Multimodal learning provides a promising solution by integrating complementary cues from various data sources. However, the heterogeneity between plant phenotypes and other modalities, such as textual descriptions, poses a significant challenge for effective fusion. To address this issue, we propose PlantIF, a multimodal feature interactive fusion model for plant disease diagnosis based on graph learning. PlantIF comprises three key components: image and text feature extractors, semantic space encoders, and a multimodal feature fusion module. Specifically, we employ pre-trained image and text feature extractors to extract visual and textual features enriched with prior knowledge of plant diseases. Semantic space encoders then map these features into both shared and modality-specific spaces, enabling the capture of cross-modal and unique semantic information. To enhance context understanding, we design a multimodal feature fusion module to process and fuse different modal semantic information, and then extract the spatial dependency between plant phenotype and text semantics through the self-attention graph convolution network. We evaluate PlantIF on a multimodal plant disease dataset with 205,007 images and 410,014 texts, achieving 96.95 % accuracy, 1.49 % higher than existing models. These results demonstrate the potential of multimodal learning in plant disease diagnosis and highlight PlantIF's value in precision agriculture. Codes are available at https://github.com/GZU-SAMLab/PlantIF.

Why it matches plant phenotyping methods植物画像とテキストを統合して病害状態を診断するモデルPlantIFを開発・評価しており、植物病害という観察可能な植物状態の推定が研究の中心である。

abstractwe propose PlantIF, a multimodal feature interactive fusion model for plant disease diagnosis based on graph learning.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicCodes are available at https://github.com/GZU-SAMLab/PlantIF .Open asset ↗GZU-SAMLab/PlantIFlines:1-28
Code / dataset availability confirmedCrossref · checked 6 Sept 2026
Published17 Oct 2025AgriEngineeringCited by 2 · OpenAlex ↗

Agri-DSSA: A Dual Self-Supervised Attention Framework for Multisource Crop Health Analysis Using Hyperspectral and Image-Based Benchmarks

Aerial / UAVField / plotMultimodalMultispectral / hyperspectralLeafWhole plant / canopy / plot / fieldClassificationObject detectionPhysiological trait estimationStress / disease detection

Recent advances in hyperspectral imaging (HSI) and multimodal deep learning have opened new opportunities for crop health analysis; however, most existing models remain limited by dataset scope, lack of interpretability, and weak cross-domain generalization. To overcome these limitations, this study introduces Agri-DSSA, a novel Dual Self-Supervised Attention (DSSA) framework that simultaneously models spectral and spatial dependencies through two complementary self-attention branches. The proposed architecture enables robust and interpretable feature learning across heterogeneous data sources, facilitating the estimation of spectral proxies of chlorophyll content, plant vigor, and disease stress indicators rather than direct physiological measurements. Experiments were performed on seven publicly available benchmark datasets encompassing diverse spectral and visual domains: three hyperspectral datasets (Indian Pines with 16 classes and 10,366 labeled samples; Pavia University with 9 classes and 42,776 samples; and Kennedy Space Center with 13 classes and 5211 samples), two plant disease datasets (PlantVillage with 54,000 labeled leaf images covering 38 diseases across 14 crop species, and the New Plant Diseases dataset with over 30,000 field images captured under natural conditions), and two chlorophyll content datasets (the Global Leaf Chlorophyll Content Dataset (GLCC), derived from MERIS and OLCI satellite data between 2003–2020, and the Leaf Chlorophyll Content Dataset for Crops, which includes paired spectrophotometric and multispectral measurements collected from multiple crop species). To ensure statistical rigor and spatial independence, a block-based spatial cross-validation scheme was employed across five independent runs with fixed random seeds. Model performance was evaluated using R2, RMSE, F1-score, AUC-ROC, and AUC-PR, each reported as mean ± standard deviation with 95% confidence intervals. Results show that Agri-DSSA consistently outperforms baseline models (PLSR, RF, 3D-CNN, and HybridSN), achieving up to R2=0.86 for chlorophyll content estimation and F1-scores above 0.95 for plant disease detection. The attention distributions highlight physiologically meaningful spectral regions (550–710 nm) associated with chlorophyll absorption, confirming the interpretability of the model’s learned representations. This study serves as a methodological foundation for UAV-based and field-deployable crop monitoring systems. By unifying hyperspectral, chlorophyll, and visual disease datasets, Agri-DSSA provides an interpretable and generalizable framework for proxy-based vegetation stress estimation. Future work will extend the model to real UAV campaigns and in-field spectrophotometric validation to achieve full agronomic reliability.

Why it matches plant phenotyping methods植物のクロロフィル含量・活力・病害ストレスを画像/ハイパースペクトルから推定する新規深層学習フレームワークを開発・評価しており、植物表現型の取得・推定が中心である。

abstractthis study introduces Agri-DSSA, a novel Dual Self-Supervised Attention (DSSA) framework that simultaneously models spectral and spatial dependencies through two complementary self-attention branches.
Reproduction assets foundThe paper's Data Availability Statement explicitly deposits the authors' Agri-DSSA implementation (the computational analysis code for the phenotyping experiments) in a public GitHub repository with a commit hash. The seven benchmark datasets are cited third-party resources rather than paper-specific deposits, so only,
Code · publicThe implementation is openly available at the GitHub repository https://github.com/ Fatema-Abdulqader/Agri-DSSA-Dual-Self-Supervised-Attention-Framework/tree/main, commit 98f3863Open asset ↗pdf-page:21 lines:1-61
Code / dataset availability confirmedOpenAlex · arXiv · checked 6 Sept 2026
Published23 Sept 2025arXiv (Cornell University)Cited by 0 · OpenAlex ↗

Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data

Field / plotMultimodalRGB / grayscaleThermalWhole plant / canopy / plot / fieldSegmentation

Accurate plant segmentation in thermal imagery remains a significant challenge for high throughput field phenotyping, particularly in outdoor environments where low contrast between plants and weeds and frequent occlusions hinder performance. To address this, we present a framework that leverages synthetic RGB imagery, a limited set of real annotations, and GAN-based cross-modality alignment to enhance semantic segmentation in thermal images. We trained models on 1,128 synthetic images containing complex mixtures of crop and weed plants in order to generate image segmentation masks for crop and weed plants. We additionally evaluated the benefit of integrating as few as five real, manually segmented field images within the training process using various sampling strategies. When combining all the synthetic images with a few labeled real images, we observed a maximum relative improvement of 22% for the weed class and 17% for the plant class compared to the full real-data baseline. Cross-modal alignment was enabled by translating RGB to thermal using CycleGAN-turbo, allowing robust template matching without calibration. Results demonstrated that combining synthetic data with limited manual annotations and cross-domain translation via generative models can significantly boost segmentation performance in complex field environments for multi-model imagery.

Why it matches plant phenotyping methods熱画像による作物・雑草の分割を対象とし、合成データ、少数の実画像、GANによるモダリティ変換を組み合わせた高スループット表現型取得手法を開発・評価している。

abstractAccurate plant segmentation in thermal imagery remains a significant challenge for high throughput field phenotyping
Reproduction assets foundThe paper states its synthetic and real phenotyping image datasets (cowpea/weed RGB and thermal imagery with segmentation masks) are publicly available through the AgML framework, with an explicit authors' URL. Helios is a general simulation tool, not a paper-specific asset.
Dataset · publicSynthetic and real datasets are available through AgML 1 1 1 https://github.com/Project-AgML/AgML [ 50 ] , a centralized framework for agricultural machine learning.Open asset ↗Project-AgML/AgMLlines:339-434
Code / dataset availability confirmedarXiv · checked 15 Sept 2026
Published16 Sept 2025arXiv

WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory

MultimodalLiDAR / point cloudWhole plant / canopy / plot / fieldClassificationSegmentation

Street trees are vital to urban livability, providing ecological and social benefits. Establishing a detailed, accurate, and dynamically updated street tree inventory has become essential for optimizing these multifunctional assets within space-constrained urban environments. Given that traditional field surveys are time-consuming and labor-intensive, automated surveys utilizing Mobile Mapping Systems (MMS) offer a more efficient solution. However, existing MMS-acquired tree datasets are limited by small-scale scene, limited annotation, or single modality, restricting their utility for comprehensive analysis. To address these limitations, we introduce WHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset. Collected across two distinct cities, WHU-STree integrates synchronized point clouds and high-resolution images, encompassing 21,007 annotated tree instances across 50 species and 2 morphological parameters. Leveraging the unique characteristics, WHU-STree concurrently supports over 10 tasks related to street tree inventory. We benchmark representative baselines for two key tasks--tree species classification and individual tree segmentation. Extensive experiments and in-depth analysis demonstrate the significant potential of multi-modal data fusion and underscore cross-domain applicability as a critical prerequisite for practical algorithm deployment. In particular, we identify key challenges and outline potential future works for fully exploiting WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV/WHU-STree.

Why it matches plant phenotyping methods樹木の個体セグメンテーションと形態パラメータを含むマルチモーダルデータセットを構築し、ベンチマークする研究であり、植物個体の状態・形態抽出手法が中心である。

abstractWHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset.
Reproduction assets foundThe paper's core asset is the WHU-STree multi-modal street tree dataset (point clouds, panoramic images, 21,007 annotated tree instances, 50 species, height/DBH), which the authors state is publicly accessible via their GitHub organization WHU-USI3DV. The Zenodo DOIs in the reference list belong to cited prior datasets
Dataset · publicticular, we identify key challenges and outline potential future works for fully exploit- ing WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV /WHU-STree. Keywords: Deep learning, Tree inventory, Individual tree segmentation, Tree species classification, Multi-modal, Mobile mapping system 1. Introduction Street trees, vital to urban ecosystems, provide ecological benefits (e.g., shade (Kumar et al., 2024), air purification (Grundstrém and Pleijel, 2014), noise reductiOpen asset ↗WHU-STreepdf-raw-page:2 lines:1-35
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published21 Aug 2025Scientific reportsCited by 5 · OpenAlex ↗

Assessment of plant diversity index in degraded desert grassland using UAV hyperspectral multimodal data and Encoder-CNN.

Aerial / UAVField / plotMultimodalMultispectral / hyperspectralWhole plant / canopy / plot / fieldClassification

The biodiversity function of the desert steppe ecosystem faces many challenges under the pressure of climate change and human activities. Accurate and efficient assessment of plant diversity is critical for guiding desert steppe restoration efforts. However, desert steppe vegetation has sparse leaves and sparse distribution. It is difficult to accurately distinguish micro-vegetation types based on a single spectrum, vegetation index or texture feature, and the resolution of satellite remote sensing cannot meet the needs of high-precision diversity assessment. To this end, this study proposed a novel method for assessing plant diversity index in degraded desert grassland based on multimodal UAV hyperspectral data and Encoder-CNN. Through experiments on different modal feature combinations, spatial spectra, vegetation indices and texture features were targeted and fused. Channel Attention Fusion (CAF) was introduced into Encoder to achieve cross-layer "soft" residual fusion, the Encoder and CNN models were fused to construct a global-local co-expression structure, and finally the quantitative calculation of the plant diversity index at the pixel level was realized. The results show that the vegetation types determined by the fusion of multimodal data and deep learning are consistent with the existing species, dominant species and sub-dominant species of the actual community, and the calculated diversity index results are also consistent with the actual situation. The use of multimodal data combining spatial spectral features with index features, combined with the Encode-CNN model, can provide the most accurate information on community composition. The overall accuracy of sparse vegetation classification can reach 90.01%, and the average accuracy can reach 85.23%, which is better than single mode or traditional 3DCNN, VIT models. This study demonstrates the application potential of UAV hyperspectral multimodal technology and deep learning in the assessment of desert steppe plant diversity, providing important technical support for ecological protection and conservation.

Why it matches plant phenotyping methodsUAVハイパースペクトルとEncoder-CNNを用いて、植物多様性指数を画素レベルで定量推定する手法を開発・評価しており、植物状態の取得・抽出が研究の中心である。

abstractthis study proposed a novel method for assessing plant diversity index in degraded desert grassland based on multimodal UAV hyperspectral data and Encoder-CNN.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産1件を確認しました。
Code · publicThe codes used in this study are available at https://github.com/15204718180/encoder-cnn.Open asset ↗15204718180/encoder-cnnpdf-page:17 lines:56-74
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published8 Aug 2025Plants (Basel, Switzerland)Cited by 6 · OpenAlex ↗

KBNet: A Language and Vision Fusion Multi-Modal Framework for Rice Disease Segmentation.

RiceMultimodalLeafSegmentationDisease symptoms / severity

High-quality disease segmentation plays a crucial role in the precise identification of rice diseases. Although the existing deep learning methods can identify the disease on rice leaves to a certain extent, these methods often face challenges in dealing with multi-scale disease spots and irregularly growing disease spots. In order to solve the challenges of rice leaf disease segmentation, we propose KBNet, a novel multi-modal framework integrating language and visual features for rice disease segmentation, leveraging the complementary strengths of CNN and Transformer architectures. Firstly, we propose the Kalman Filter Enhanced Kolmogorov-Arnold Networks (KF-KAN) module, which combines the modeling ability of KANs for nonlinear features and the dynamic update mechanism of the Kalman filter to achieve accurate extraction and fusion of multi-scale lesion information. Secondly, we introduce the Boundary-Constrained Physical-Information Neural Network (BC-PINN) module, which embeds the physical priors, such as the growth law of the lesion, into the loss function to strengthen the modeling of irregular lesions. At the same time, through the boundary punishment mechanism, the accuracy of edge segmentation is further improved and the overall segmentation effect is optimized. The experimental results show that the KBNet framework demonstrates solid performance in handling complex and diverse rice disease segmentation tasks and provides key technical support for disease identification, prevention, and control in intelligent agriculture. This method has good popularization value and broad application potential in agricultural intelligent monitoring and management.

Why it matches plant phenotyping methodsイネ葉の病斑という植物の病態を画像からセグメンテーションする新規手法を開発しており、病害識別のための表現抽出・境界推定が研究の中心である。

abstractwe propose KBNet, a novel multi-modal framework integrating language and visual features for rice disease segmentation
Reproduction assets foundThe paper's rice disease segmentation images (1550 leaf images of Blast, Bacterial Blight, and Tungro) were sourced from the public PlantVillage dataset on Kaggle, which is publicly available at the allowed URL. However, the paper-specific derived assets — the authors' pixel-level Labelme segmentation masks, text描述, 4:
Dataset · publici: 10.1007/s10462-024-10717-2. 26. Li A., Jiao L., Zhu H., Li L., Liu F. Multitask Semantic Boundary Awareness Network for Remote Sensing Image Segmentation. IEEE Trans. Geosci. Remote Sens. 2022;60:1–14. doi: 10.1109/TGRS.2021.3050885. 27. Kaggle, PlantVillage Dataset. 2019. [(accessed on 19 September 2022)]. Available online: https://www.kaggle.com/datasets/abdallahalidev/plantvillage-dataset . 28. Li Z., Li Y., Li Q., Wang P., Guo D., Lu L., Jin D., Zhang Y., Hong Q. LViT: Language Meets Vision Transformer in Medical Image Segmentation. IEEE Trans. Med. Imaging. 2024;43:96–107. doi: 10.1109/TMI.2023.3291719. 29. Liu Z., Wang Y., Vaidya S., Ruehle F., Halverson J., Soljacic M., Hou T.Y., TOpen asset ↗Kaggle · plantvillage-datasetlines:409-431
Code / dataset availability confirmedCrossref · Europe PMC · checked 6 Sept 2026
Published30 Jun 2025SensorsCited by 14 · OpenAlex ↗

Cross-Modal Data Fusion via Vision-Language Model for Crop Disease Recognition.

SoybeanMultimodalLeafClassificationStress / disease detectionDisease symptoms / severity

Crop diseases pose a significant threat to agricultural productivity and global food security. Timely and accurate disease identification is crucial for improving crop yield and quality. While most existing deep learning-based methods focus primarily on image datasets for disease recognition, they often overlook the complementary role of textual features in enhancing visual understanding. To address this problem, we proposed a cross-modal data fusion via a vision-language model for crop disease recognition. Our approach leverages the Zhipu.ai multi-model to generate comprehensive textual descriptions of crop leaf diseases, including global description, local lesion description, and color-texture description. These descriptions are encoded into feature vectors, while an image encoder extracts image features. A cross-attention mechanism then iteratively fuses multimodal features across multiple layers, and a classification prediction module generates classification probabilities. Extensive experiments on the Soybean Disease, AI Challenge 2018, and PlantVillage datasets demonstrate that our method outperforms state-of-the-art image-only approaches with higher accuracy and fewer parameters. Specifically, with only 1.14M model parameters, our model achieves a 98.74%, 87.64% and 99.08% recognition accuracy on the three datasets, respectively. The results highlight the effectiveness of cross-modal learning in leveraging both visual and textual cues for precise and efficient disease recognition, offering a scalable solution for crop disease recognition.

Why it matches plant phenotyping methods作物葉の病徴を画像・テキストから認識するマルチモーダル手法の開発と評価が研究の中心であり、植物の病害状態を直接推定している。

abstractwe proposed a cross-modal data fusion via a vision-language model for crop disease recognition.
Reproduction assets foundThe paper's Data Availability Statement explicitly links the public image datasets used for its crop disease recognition experiments: the Soybean Disease dataset (Dryad DOI) and the PlantVillage dataset (Kaggle). These are the phenotyping image inputs directly used in this study. The AI Challenge 2018 dataset is also公开
Dataset · publicThe soybean dataset is available at https://doi.org/10.5061/dryad.41ns1rnj3 (accessed on 1 April 2025).Open asset ↗Dryad · 10.5061/dryad.41ns1rnj3pdf-page:12 lines:1-58
Dataset · publicplantvillage dataset is available at https://www.kaggle.com/datasets/abdallahalidev/plantvillage-Open asset ↗Kagglepdf-page:12 lines:1-58
Code / dataset availability confirmedOpenAlex · Crossref · Europe PMC · checked 6 Sept 2026
Published27 Jun 2025Science AdvancesCited by 43 · OpenAlex ↗

A machine-learning-powered spectral-dominant multimodal soft wearable system for long-term and early-stage diagnosis of plant stresses.

TomatoGreenhouseMultimodalMultispectral / hyperspectralLeafStress / disease detectionStress response / tolerancePlant / canopy temperature

Addressing the global malnutrition crisis requires precise and timely diagnostics of plant stresses to enhance the quality and yield of nutrient-rich crops, such as tomatoes. Soft wearable sensors offer a promising approach by continuously monitoring plant physiology. However, challenges remain in identifying direct physiological indicators of plant stresses, hindering the development of accurate diagnostic models for predicting symptom progression. Here, we introduce a machine-learning-powered spectral-dominant multimodal soft wearable system (MapS-Wear) for precise, long-term, and early-stage diagnosis of stresses in tomatoes. MapS-Wear continuously tracks leaf surrounding temperature, humidity, and unique in-situ transmission spectra, which are critical stress-related indicators. The machine learning framework processes these multimodal data to predict gradual stress progression and diagnose nutrient deficiencies in plants over 10 days earlier than conventional computer vision methods. Moreover, MapS-Wears enables portable and large-scale screening of grafted tomato varieties in greenhouses, accelerating the identification of compatible grafting combinations. This demonstration highlights the potential for high-throughput plant phenotyping and yield improvement.

Why it matches plant phenotyping methods植物ストレスの生理状態を連続センシングし、機械学習で早期診断・進行予測するウェアラブル計測システムが研究の中心であり、植物フェノタイピング手法として明確に該当する。

abstractHere, we introduce a machine-learning-powered spectral-dominant multimodal soft wearable system (MapS-Wear) for precise, long-term, and early-stage diagnosis of stresses in tomatoes.
Reproduction assets foundThe paper's Data and materials availability statement explicitly deposits the tomato leaf photos, transmission spectral data, and ML algorithms on Zenodo, matching an allowed URL.
Dataset · publicThe photos of tomato leaves in different health statuses, the transmission spectral data of these leaves, and the ML algorithms are openly available on Zenodo ( https://zenodo.org/doi/10.5281/zenodo.15192884 ).Open asset ↗Zenodo · 10.5281/zenodo.15192884lines:129-274
Code / dataset availability confirmedOpenAlex · checked 6 Sept 2026
Published6 Dec 2024Plant CommunicationsCited by 16 · OpenAlex ↗

Phenomics-assisted genetic dissection and molecular design of drought resistance in rice

RiceField / plotMultimodalPanicle / ear / spikeLeafRootSeed / grainGrowth / time-series analysisBiomass / plant weightLeaf traits

Dissecting the drought resistance (DR) mechanism and designing drought-resistant rice varieties are promising strategies to address the challenge of climate change. Here, we selected a typical drought-avoidant (DA) variety IRAT109 and drought-tolerant (DT) variety Hanhui15 as the parents to develop a stable recombinant inbred line (RIL) population (F 8 , 1,262 lines). The de novo assembled genomes of both parents were released. Through re-sequencing of the RIL population, a set of 1,189,216 reliable SNPs were obtained and used for constructing a dense genetic map. Using both aboveground and underground phenomic platforms and multimodal cameras, we captured 139,040 image-based traits (i-traits) of whole plant’s phenotypes in response to drought stress throughout entire rice growth period and identified 32,586 drought-responsive quantitative trait loci (QTLs) including 2,097 unique QTLs. The QTLs related to panicle i-traits occurred on the middle of chromosome 8 over 600 times, while the QTLs related to leaf i-traits on the 5’ end of chromosome 3 over 800 times, indicating potential effect of these QTLs on plant phenotypes. We chose three candidate genes ( OsMADS50, OsGhd8, OsSAUR11 ) related to leaf, panicle, and root traits respectively and verified their functions in resisting drought. Gene OsMADS50 was found to negatively regulate DR by modulating leaf dehydration, grain size, and root downward growth. Furthermore, a total of 18 and 21 composite QTLs significantly related to grain weight and plant biomass were screened from 597 lines in RIL population under drought conditions in field experiments, and composite QTL region was highly overlapped (76.9%) with known DR gene region. Based on three candidate DR genes, we proposed the haplotype design suitable for different environments and breeding objectives. This study provides a valuable reference for multi-modal and time-series phenomic analyses, deciphers the genetic mechanism of DA and DT rice varieties, and offers a molecular navigation map for breeding DR variety.

Why it matches plant phenotyping methods地下・地上フェノミックプラットフォームとマルチモーダルカメラで全生育期間の画像形質を大量取得しており、フェノタイピング手法の適用と技術的ワークフローが研究の中核です。

abstractUsing both aboveground and underground phenomic platforms and multimodal cameras, we captured 139,040 image-based traits (i-traits) of whole plant’s phenotypes in response to drought stress throughout entire rice growth period
Reproduction assets foundThe paper's phenome data (aboveground and belowground rice images/i-traits) and the authors' data-handling code and deep-learning model are explicitly deposited at public URLs listed in the Data Availability Statement. Genome data (riceome.hzau.edu.cn) is molecular omics and excluded.
Code · publicAll the phenome data and core data-handling code have been deposited online.Open asset ↗lines:140-175
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 7 Sept 2026
Published15 Nov 2024PlantsCited by 11 · OpenAlex ↗

Multimodal Data Fusion for Precise Lettuce Phenotype Estimation Using Deep Learning Algorithms

LettuceMultimodalRGB / grayscaleLeafWhole plant / canopy / plot / fieldMorphology / geometry measurementObject detectionSegmentationArchitecture / morphology / geometryBiomass / plant weight

Effective lettuce cultivation requires precise monitoring of growth characteristics, quality assessment, and optimal harvest timing. In a recent study, a deep learning model based on multimodal data fusion was developed to estimate lettuce phenotypic traits accurately. A dual-modal network combining RGB and depth images was designed using an open lettuce dataset. The network incorporated both a feature correction module and a feature fusion module, significantly enhancing the performance in object detection, segmentation, and trait estimation. The model demonstrated high accuracy in estimating key traits, including fresh weight (fw), dry weight (dw), plant height (h), canopy diameter (d), and leaf area (la), achieving an R2 of 0.9732 for fresh weight. Robustness and accuracy were further validated through 5-fold cross-validation, offering a promising approach for future crop phenotyping.

Why it matches plant phenotyping methodsRGB・深度画像を融合した深層学習によるレタス形質推定手法を開発し、交差検証で性能評価しており、フェノタイピング手法が中心である。

abstracta deep learning model based on multimodal data fusion was developed to estimate lettuce phenotypic traits accurately
Reproduction assets foundThe paper's RGB-D lettuce images and trait measurements come from the publicly available Third Autonomous Greenhouse Challenge dataset deposited at 4TU.ResearchData, with an explicit availability statement and URL matching an allowed URL. No author analysis code or trained model is disclosed.
Dataset · publicThis study used the Third Autonomous Greenhouse Challenge: Online Challenge Lettuce Images dataset publicly available at 4TU.ResearchData [ 36 ].Open asset ↗4TU.ResearchDatalines:819-832
Code / dataset availability confirmedCrossref · checked 15 Sept 2026
Published7 Sept 2024AgronomyCited by 8 · OpenAlex ↗

Crop Growth Analysis Using Automatic Annotations and Transfer Learning in Multi-Date Aerial Images and Ortho-Mosaics

Brassica vegetablesAerial / UAVMultimodalWhole plant / canopy / plot / fieldAnnotation / quality controlObject detectionSegmentationGrowth / time-series analysisGrowth / development / phenologyYield / yield components

Growth monitoring of crops is a crucial aspect of precision agriculture, essential for optimal yield prediction and resource allocation. Traditional crop growth monitoring methods are labor-intensive and prone to errors. This study introduces an automated segmentation pipeline utilizing multi-date aerial images and ortho-mosaics to monitor the growth of cauliflower crops (Brassica Oleracea var. Botrytis) using an object-based image analysis approach. The methodology employs YOLOv8, a Grounding Detection Transformer with Improved Denoising Anchor Boxes (DINO), and the Segment Anything Model (SAM) for automatic annotation and segmentation. The YOLOv8 model was trained using aerial image datasets, which then facilitated the training of the Grounded Segment Anything Model framework. This approach generated automatic annotations and segmentation masks, classifying crop rows for temporal monitoring and growth estimation. The study’s findings utilized a multi-modal monitoring approach to highlight the efficiency of this automated system in providing accurate crop growth analysis, promoting informed decision-making in crop management and sustainable agricultural practices. The results indicate consistent and comparable growth patterns between aerial images and ortho-mosaics, with significant periods of rapid expansion and minor fluctuations over time. The results also indicated a correlation between the time and method of observation which paves a future possibility of integration of such techniques aimed at increasing the accuracy in crop growth monitoring based on automatically derived temporal crop row segmentation masks.

Why it matches plant phenotyping methods航空画像・オルソモザイクから作物列を自動セグメンテーションし、時系列の生育・成長を推定する画像解析パイプラインが研究の中心であり、植物形質の取得手法として適格。

abstractThis study introduces an automated segmentation pipeline utilizing multi-date aerial images and ortho-mosaics to monitor the growth of cauliflower crops
Reproduction assets foundThe paper's Data Availability Statement points to the authors' public Mendeley Data repository (GobhiSet, DOI 10.17632/dcjjcwc5dh.4), which contains the raw, manually, and automatically annotated RGB aerial images and ortho-mosaics of cauliflower used for the YOLOv8x-seg and Grounded SAM training and growth analysis in
Dataset · publicon of the manuscript. Funding: This research received no external funding. Data Availability Statement: No new data was created. However, the data that were used to perform this research can be found in the article published at https://doi.org/10.1016/j.dib.2024.110506 and available in the repository DOI: 10.17632/dcjjcwc5dh.4 (https://data.mendeley.com/drafts/dcjjcwc5dh).Conflicts of Interest: The authors declare no conflicts of interest. References 1. Di, L.; Ustundag, B. Crop Growth Modeling and Yield Forecasting. In Agro-Geoinformatics; Springer: Cham, Switzerland, 2021. [CrossRef] 2. Mithen, S.; Jenkins, E.; Jamjoum, K.; Nuimat, S.; Nortcliff, S.; Finlayson, B. Experimental crop growingOpen asset ↗data.mendeley.com · 10.17632/dcjjcwc5dh.4pdf-raw-page:17 lines:1-52
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published13 Aug 2024Plant phenomics (Washington, D.C.)Cited by 26 · OpenAlex ↗

A Multi-Modal Open Object Detection Model for Tomato Leaf Diseases with Strong Generalization Performance Using PDC-VLD.

TomatoMultimodalLeafObject detectionDisease symptoms / severity

Precise disease detection is crucial in modern precision agriculture, especially in ensuring the health of tomato crops and enhancing agricultural productivity and product quality. Although most existing disease detection methods have helped growers identify tomato leaf diseases to some extent, these methods typically target fixed categories. When faced with new diseases, extensive and costly manual annotation is required to retrain the dataset. To overcome these limitations, this study proposes a multimodal model PDC-VLD based on the open-vocabulary object detection (OVD) technology within the VLDet framework, which can accurately identify new tomato leaf diseases without manual annotation by using only image-text pairs. First, we developed a progressive visual transformer-convolutional pyramid module (PVT-C) that effectively extracts tomato leaf disease features and optimizes anchor box positioning using the self-supervised learning algorithm DINO, suppressing interference from irrelevant backgrounds. Then, a context feature guided module (CFG) was adopted to address the low adaptability and recognition accuracy of the model in data-scarce environments. To validate the model's effectiveness, we constructed a tomato leaf disease image dataset containing 4 base classes and 2 new categories. Experimental results show that the PDC-VLD model achieved 61.2% on the main evaluation metric mAPnovel50 , and 56.4% on mAPnovel75 , 87.7% on mAPbase50 , 81.0% on mAPall50 , and 45.5% on average recall, outperforming existing OVD models. Our research provides an innovative solution for efficiently and accurately detecting new diseases, substantially reducing the need for manual annotation, and offering critical technical support and practical reference for agricultural workers.

Why it matches plant phenotyping methodsトマト葉の病害状態を画像から検出するマルチモーダルモデルを開発し、データセット構築と性能評価まで行っており、植物フェノタイピング手法が研究の中心である。

abstractthis study proposes a multimodal model PDC-VLD based on the open-vocabulary object detection (OVD) technology
Reproduction assets foundThe paper's Data Availability statement says all datasets used were uploaded to the authors' public GitHub repository (PDC-VLD), and the paper also uses the public Kaggle PlantVillage dataset as a plant image source. Bespoke portions of the dataset require contacting the corresponding author.
Dataset · publicAll datasets that were used and analyzed in this study have been uploaded to the website https://github.com/ZhouGuoXiong/PDC-VLD . Furthermore, for access to all bespoke datasets used in this study (comprising a total of 6,923 images and 13,864 texts), please contact the corresponding author.Open asset ↗ZhouGuoXiong/PDC-VLDlines:797-797
Code / dataset availability confirmedCrossref · checked 7 Sept 2026
Published23 Jul 2024Remote SensingCited by 10 · OpenAlex ↗

Monitoring Cover Crop Biomass in Southern Brazil Using Combined PlanetScope and Sentinel-1 SAR Data

RyeField / plotMultimodalMultispectral / hyperspectralWhole plant / canopy / plot / fieldYield / biomass estimationBiomass / plant weight

Precision agriculture integrates multiple sensors and data types to support farmers with informed decision-making tools throughout crop cycles. This study evaluated Aboveground Biomass (AGB) estimates of Rye using attributes derived from PlanetScope (PS) optical, Sentinel-1 Synthetic Aperture Radar (SAR), and hybrid (optical plus SAR) datasets. Optical attributes encompassed surface reflectance from PS’s blue, green, red, and near-infrared (NIR) bands, alongside the Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EVI). Sentinel-1 SAR attributes included the C-band Synthetic Aperture Radar Ground Range Detected, VV and HH polarizations, and both Ratio and Polarization (Pol) indices. Ground reference AGB data for Rye (Secale cereal L.) were collected from 50 samples and four dates at a farm located in southern Brazil, aligning with image acquisition dates. Multiple linear regression models were trained and validated. AGB was estimated based on individual (optical PS or Sentinel-1 SAR) and combined datasets (optical plus SAR). This process was repeated 100 times, and variable importance was extracted. Results revealed improved Rye AGB estimates with integrated optical and SAR data. Optical vegetation indices displayed higher correlation coefficients (r) for AGB estimation (r = +0.67 for both EVI and NDVI) compared to SAR attributes like VV, Ratio, and polarization (r ranging from −0.52 to −0.58). However, the hybrid regression model enhanced AGB estimation (R2 = 0.62, p

Why it matches plant phenotyping methods光学・SARセンサーデータを統合してライムギの地上部バイオマスを推定し、回帰モデルを訓練・検証しているため、植物形質の取得・推定手法が研究の中心です。

abstractThis study evaluated Aboveground Biomass (AGB) estimates of Rye using attributes derived from PlanetScope (PS) optical, Sentinel-1 Synthetic Aperture Radar (SAR), and hybrid (optical plus SAR) datasets.
Reproduction assets foundThe paper's Sentinel-1 SAR processing workflow is publicly shared as a Google Earth Engine JavaScript script with an explicit availability statement and URL. The field AGB measurements and PlanetScope data are not public (available only on reasonable request from the corresponding author).
Code · publicData Availability Statement: The Google Earth Engine script to process Sentinel-1 data is available at [https://code.earthengine.google.com/219fc3c05b8ae8132ae2d758ccba3d1e?noload=true] (accessed on 12 July 2024).Open asset ↗pdf-page:17 lines:1-59
Code / dataset availability confirmedOpenAlex · arXiv · checked 15 Sept 2026
Published3 Jul 2024arXiv (Cornell University)Cited by 1 · OpenAlex ↗

3D Multimodal Image Registration for Plant Phenotyping

MultimodalRGB-D / ToFWhole plant / canopy / plot / fieldImage / point-cloud registration

The use of multiple camera technologies in a combined multimodal monitoring system for plant phenotyping offers promising benefits. Compared to configurations that only utilize a single camera technology, cross-modal patterns can be recorded that allow a more comprehensive assessment of plant phenotypes. However, the effective utilization of cross-modal patterns is dependent on precise image registration to achieve pixel-accurate alignment, a challenge often complicated by parallax and occlusion effects inherent in plant canopy imaging. In this study, we propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process. By leveraging depth data, our method mitigates parallax effects and thus facilitates more accurate pixel alignment across camera modalities. Additionally, we introduce an automated mechanism to identify and differentiate different types of occlusions, thereby minimizing the introduction of registration errors. To evaluate the efficacy of our approach, we conduct experiments on a diverse image dataset comprising six distinct plant species with varying leaf geometries. Our results demonstrate the robustness of the proposed registration algorithm, showcasing its ability to achieve accurate alignment across different plant types and camera compositions. Compared to previous methods it is not reliant on detecting plant specific image features and can thereby be utilized for a wide variety of applications in plant sciences. The registration approach principally scales to arbitrary numbers of cameras with different resolutions and wavelengths. Overall, our study contributes to advancing the field of plant phenotyping by offering a robust and reliable solution for multimodal image registration.

Why it matches plant phenotyping methods植物フェノタイピングのためのマルチモーダル3D画像登録手法を開発し、複数植物種の画像データセットで性能評価しているため、方法が研究の中心である。

abstractIn this study, we propose a novel multimodal 3D image registration method that addresses these challenges by integrating depth information from a time-of-flight camera into the registration process.
Reproduction assets foundThe paper's multimodal plant image dataset (six plant species recorded with the RGBD/thermal/hyperspectral setup) is publicly available on the authors' GitHub repository, which is an allowed URL.
Dataset · publiclity of our registration algorithm across diverse scenarios, we recorded a dataset comprising images of six distinct plant species. This was done to encompass a wide variety of leaf and canopy structures, thus offering a representative sample for evaluation purposes. The recorded dataset can be found on the project github page: https://github.com/eric-stumpe/Plant3DImageReg . The chosen plant species are as follows: 1. Grapevine ( Vitis vinifera ) 2. Leopard lily ( Dieffenbachia ) 3.Open asset ↗eric-stumpe/Plant3DImageReglines:287-310
Code / dataset availability confirmedEurope PMC · checked 14 Sept 2026
Published3 May 2024Data in briefCited by 13 · OpenAlex ↗

EscaYard: Precision viticulture multimodal dataset of vineyards affected by Esca disease consisting of geotagged smartphone images, phytosanitary status, UAV 3D point clouds and Orthomosaics.

GrapevineAerial / UAVField / plotMultimodalLiDAR / point cloudMultispectral / hyperspectralFruitLeafWhole plant / canopy / plot / fieldDisease symptoms / severity

The "EscaYard" dataset comprises multimodal data collected from vineyards to support agricultural research, specifically focusing on vine health and productivity. Data collection involved two primary methods: (1) unmanned aerial vehicle (UAV) for capturing multispectral images and 3D point clouds, and (2) smartphones for detailed ground-level photography. The UAV used was DJI Matrice 210 V2 RTK, equipped with a Micasense Altum sensor, flying at 30 m above ground level to ensure detailed coverage. Ground-level data were collected using smartphones (iPhone X and Xiaomi Poco X3 Pro), which provided high-resolution images of individual plants. These images were geotagged, enabling location mapping, and included data on the phytosanitary status and number of grape clusters per plant. Additionally, the dataset contains RTK GNSS data, offering high-precision location information for each vine, enhancing the dataset's value for spatial analysis. Moreover, the dataset is structured to support various research applications, including agronomy, remote sensing, and machine learning. It is particularly suited for studying disease detection, yield estimation, and vineyard management strategies. The high-resolution and multispectral nature of the data allows for a detailed analysis of vineyard conditions. Potential reuse of the dataset spans multiple disciplines, enabling studies on environmental monitoring, geographic information systems (GIS), and precision agriculture. Its comprehensive nature makes it a valuable resource for developing and testing algorithms for disease classification, yield prediction, and plant phenotyping. For instance, the images of bunches and grape leaves can be used to train object detection algorithms for accurate disease detection and consequent precise spraying. Moreover, yield prediction algorithms can be trained by extracting the phenotypic traits of the grape bunches. The "EscaYard" dataset provides a foundation for advancing research in sustainable farming practices, optimising crop health, and improving productivity through precise agricultural technologies.

Why it matches plant phenotyping methodsブドウの病徴・生産性・房形質を対象とするマルチモーダル画像/UAVデータセットであり、植物フェノタイピングや病害・収量推定アルゴリズムの開発と評価を主目的としているため。

abstractThe "EscaYard" dataset provides a foundation for advancing research in sustainable farming practices, optimising crop health, and improving productivity through precise agricultural technologies.
Reproduction assets foundThe paper is a Data in Brief article describing the EscaYard dataset, publicly deposited on Zenodo with explicit DOI and direct URL. The dataset contains the paper's own phenotyping measurements (geotagged smartphone images, phytosanitary status, grape cluster counts, UAV orthomosaics, 3D point clouds, RTK GNSS trunk-­
Dataset · publics City/Town/Region: Tomiño, Pontevedra, Galicia Country: Spain Coordinates: Vineyard B7, X: 517183.8, Y: 4645072.8; Vineyard B9, X: 516987.8, Y: 4644823.7 (ETRS89 / UTM zone 29N, EPSG:25829). Data accessibility Repository name: Zenodo Data identification number: https://zenodo.org/doi/10.5281/zenodo.10362567 Direct URL to data: https://zenodo.org/records/10362567 1. Value of the Data • The dataset offers a unique combination of multimodal data, including geotagged smartphone images, UAV orthomosaics, 3D point clouds, and precise geolocation data, enabling a multifaceted analysis of vineyard health and productivity. •Open asset ↗Zenodo · 10.5281/zenodo.10362567lines:1-49
Code / dataset availability confirmedOpenAlex · Europe PMC · Crossref · checked 15 Sept 2026
Published31 Oct 2023Frontiers in Plant ScienceCited by 7 · OpenAlex ↗

OSC-CO2: coattention and cosegmentation framework for plant state change with multiple features

Chlorophyll fluorescenceMultimodalWhole plant / canopy / plot / fieldSegmentationGrowth / time-series analysisGrowth / development / phenology

Cosegmentation and coattention are extensions of traditional segmentation methods aimed at detecting a common object (or objects) in a group of images. Current cosegmentation and coattention methods are ineffective for objects, such as plants, that change their morphological state while being captured in different modalities and views. The Object State Change using Coattention-Cosegmentation (OSC-CO2) is an end-to-end unsupervised deep-learning framework that enhances traditional segmentation techniques, processing, analyzing, selecting, and combining suitable segmentation results that may contain most of our target object’s pixels, and then displaying a final segmented image. The framework leverages coattention-based convolutional neural networks (CNNs) and cosegmentation-based dense Conditional Random Fields (CRFs) to address segmentation accuracy in high-dimensional plant imagery with evolving plant objects. The efficacy of OSC-CO2 is demonstrated using plant growth sequences imaged with infrared, visible, and fluorescence cameras in multiple views using a remote sensing, high-throughput phenotyping platform, and is evaluated using Jaccard index and precision measures. We also introduce CosegPP+, a dataset that is structured and can provide quantitative information on the efficacy of our framework. Results show that OSC-CO2 out performed state-of-the art segmentation and cosegmentation methods by improving segementation accuracy by 3% to 45%.

Why it matches plant phenotyping methods植物画像から成長状態を抽出する画像セグメンテーション手法を開発し、ハイスループット表現型解析プラットフォームで評価しているため、方法が研究の中心である。

abstractThe Object State Change using Coattention-Cosegmentation (OSC-CO2) is an end-to-end unsupervised deep-learning framework
Reproduction assets foundThe paper's authors publicly released both their analysis code (OSC-CO2 framework on GitHub) and the paper-specific plant phenotyping image dataset (CosegPP+, a VSTEM plant imagery dataset from the UNL LemnaTec platform) with explicit availability statements and URLs.
Code · publicthe object’s (plant’s) shape, orientation, and size at a specific point in time. OSC-CO 2 is designed to process datasets that contain a variety of features, such as perspective (V), species (S), temporality (T), environmental conditions (E) and modality (M) (VSTEM) ( Figure 1 ). The code for OSC-CO 2 is publicly available at: https://github.com/rubiquinones/OSC-CO2 . Figure 1 A preview of a VSTEM Dataset. This work will use the CosegPP dataset ( Quiñones et al., 2021 ) and modify it as CosegPP+ and categorize it as a VSTEM dataset for our problem definition. The first row shows the growth sequence of a Buckwheat plant from 3 rd July 2019 to 27 th July 2019. The second row shows the threeOpen asset ↗rubiquinones/OSC-CO2lines:37-49
Dataset · publicasets through segmentation using Otsu’s method ( Otsu, 1979 ) and cosegmentation using Subdiscover ( Meng et al., 2016 ). These two methods were chosen since ( Quiñones et al., 2021 ) defined these as the top methods for being able to segment some of the challenging features of computer vision. CosegPP+ is publicly available at https://doi.org/10.5281/zenodo.6863013 . We replaced the original images with the outputs generated by Otsu’s method and Subdiscover. Meaning that each time point i will have at most a binary images where a is the number of algorithms (i.e., Otsu’s method and Subdiscover) used. Some groups do not contain Subdiscover binary masks due to the method’s limitation in notOpen asset ↗10.5281/zenodo.6863013lines:350-411
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 14 Sept 2026
Published9 Jan 2023PlantsCited by 10 · OpenAlex ↗

An Open-Source Package for Thermal and Multispectral Image Analysis for Plants in Glasshouse.

GreenhouseMultimodalMultispectral / hyperspectralThermalWhole plant / canopy / plot / fieldCalibration / preprocessingStress / disease detection

Advanced plant phenotyping techniques to measure biophysical traits of crops are helping to deliver improved crop varieties faster. Phenotyping of plants using different sensors for image acquisition and its analysis with novel computational algorithms are increasingly being adapted to measure plant traits. Thermal and multispectral imagery provides novel opportunities to reliably phenotype crop genotypes tested for biotic and abiotic stresses under glasshouse conditions. However, optimization for image acquisition, pre-processing, and analysis is required to correct for optical distortion, image co-registration, radiometric rescaling, and illumination correction. This study provides a computational pipeline that optimizes these issues and synchronizes image acquisition from thermal and multispectral sensors. The image processing pipeline provides a processed stacked image comprising RGB, green, red, NIR, red edge, and thermal, containing only the pixels present in the object of interest, e.g., plant canopy. These multimodal outputs in thermal and multispectral imageries of the plants can be compared and analysed mutually to provide complementary insights and develop vegetative indices effectively. This study offers digital platform and analytics to monitor early symptoms of biotic and abiotic stresses and to screen a large number of genotypes for improved growth and productivity. The pipeline is packaged as open source and is hosted online so that it can be utilized by researchers working with similar sensors for crop phenotyping.

Why it matches plant phenotyping methods植物の熱画像・マルチスペクトル画像を用いた表現型取得と解析のためのオープンソース計算パイプラインを開発しており、画像補正・共登録・解析が中心的な方法論的貢献である。

abstractThis study provides a computational pipeline that optimizes these issues and synchronizes image acquisition from thermal and multispectral sensors.
Reproduction assets found保存済みの本文根拠を更新済みルールで再検証し、公開資産2件を確認しました。
Code · publicAll codes were written in MATLAB to produce a library package which is available at https://github.com/SmartSense-iHub/Thermal-and-Multispectral-Image-Analysis-Processing-Pipeline.git (accessed on 12 November 2022).Open asset ↗SmartSense-iHub/Thermal-and-Multispectral-Image-Analysis-Processing-Pipelinelines:35-43
Dataset · publicThe data is freely shared in google drive and can be accessed from the following link. https://drive.google.com/file/d/1VSqRu5CUZhyd3MF23kdRjqrtRke7sbJU/view?usp=share_link .Open asset ↗lines:91-241
Code / dataset availability confirmedEurope PMC · checked 15 Sept 2026
Published5 Jan 2023Frontiers in plant scienceCited by 21 · OpenAlex ↗

Visual question answering model for fruit tree disease decision-making based on multimodal deep learning.

MultimodalWhole plant / canopy / plot / fieldClassificationStress / disease detectionDisease symptoms / severity

Visual Question Answering (VQA) about diseases is an essential feature of intelligent management in smart agriculture. Currently, research on fruit tree diseases using deep learning mainly uses single-source data information, such as visible images or spectral data, yielding classification and identification results that cannot be directly used in practical agricultural decision-making. In this study, a VQA model for fruit tree diseases based on multimodal feature fusion was designed. Fusing images and Q&A knowledge of disease management, the model obtains the decision-making answer by querying questions about fruit tree disease images to find relevant disease image regions. The main contributions of this study were as follows: (1) a multimodal bilinear factorized pooling model using Tucker decomposition was proposed to fuse the image features with question features: (2) a deep modular co-attention architecture was explored to simultaneously learn the image and question attention to obtain richer graphical features and interactivity. The experiments showed that the proposed unified model combining the bilinear model and co-attentive learning in a new network architecture obtained 86.36% accuracy in decision-making under the condition of limited data (8,450 images and 4,560k Q&A pairs of data), outperforming existing multimodal methods. The data augmentation is adopted on the training set to avoid overfitting. Ten runs of 10-fold cross-validation are used to report the unbiased performance. The proposed multimodal fusion model achieved friendly interaction and fine-grained identification and decision-making performance. Thus, the model can be widely deployed in intelligent agriculture.

Why it matches plant phenotyping methods果樹病害画像から病害状態を推定するマルチモーダルVQAモデルの開発と性能検証が研究の中心であり、植物病害表現型の画像ベース推定に該当する。

abstractIn this study, a VQA model for fruit tree diseases based on multimodal feature fusion was designed.
Reproduction assets foundThe authors explicitly provide a public GitHub repository containing the paper-specific VQA fruit tree disease dataset (8,450 images with Q&A triplets) and the model implementation.
Code · publica regularization to save model parameters after each epoch to prevent overfitting. To evaluate the model, we chose the best epoch based on the accuracy of the validation set. The proposed method was trained using the PyTorch library, and the experiments were run on Nvidia GTX3090 Ti 32GB GPU. The implementation is available at https://github.com/guoyaqi1/vqa_Fruit-tree-disease . After tuning all model parameters by training, we trained the model once on all available data (training set + validation set). Finally, we evaluated the test set to obtain the evaluation results of the model. The main hyperparameter setting is shown in Table 4 . Most of the values are set by trial-and-error method, Open asset ↗guoyaqi1/vqa_Fruit-tree-diseaselines:286-468
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 8 Sept 2026
Published25 Aug 2022Frontiers in plant scienceCited by 43 · OpenAlex ↗

Automatic monitoring of lettuce fresh weight by multi-modal fusion based deep learning

LettuceGrowth chamberMultimodalRGB-D / ToFLeafRootStem / branchWhole plant / canopy / plot / fieldSegmentationYield / biomass estimation

Fresh weight is a widely used growth indicator for quantifying crop growth. Traditional fresh weight measurement methods are time-consuming, laborious, and destructive. Non-destructive measurement of crop fresh weight is urgently needed in plant factories with high environment controllability. In this study, we proposed a multi-modal fusion based deep learning model for automatic estimation of lettuce shoot fresh weight by utilizing RGB-D images. The model combined geometric traits from empirical feature extraction and deep neural features from CNN. A lettuce leaf segmentation network based on U-Net was trained for extracting leaf boundary and geometric traits. A multi-branch regression network was performed to estimate fresh weight by fusing color, depth, and geometric features. The leaf segmentation model reported a reliable performance with a mIoU of 0.982 and an accuracy of 0.998. A total of 10 geometric traits were defined to describe the structure of the lettuce canopy from segmented images. The fresh weight estimation results showed that the proposed multi-modal fusion model significantly improved the accuracy of lettuce shoot fresh weight in different growth periods compared with baseline models. The model yielded a root mean square error (RMSE) of 25.3 g and a coefficient of determination ( R 2 ) of 0.938 over the entire lettuce growth period. The experiment results demonstrated that the multi-modal fusion method could improve the fresh weight estimation performance by leveraging the advantages of empirical geometric traits and deep neural features simultaneously.

Why it matches plant phenotyping methodsRGB-D画像からレタスの生体重を非破壊推定する画像解析・深層学習手法の開発が研究の中心であり、植物表現型取得法に該当する。

abstractA lettuce leaf segmentation network based on U-Net was trained for extracting leaf boundary and geometric traits.
Reproduction assets foundThe paper's phenotyping inputs (top-view RGB and aligned depth images of 388 lettuces with destructively measured traits) come from the publicly available 3rd Autonomous Greenhouse Challenge Online Challenge Lettuce Images dataset, with an explicit public URL in the data availability statement. No author analysis code,
Dataset · publicPublicly available datasets were analyzed in this study. This data can be found here: https://data.4tu.nl/articles/dataset/3rd_Autonomous_Greenhouse_Challenge_Online_Challenge_Lettuce_Images/15023088 .Open asset ↗data.4tu.nl · 15023088lines:657-691
Code / dataset availability confirmedbioRxiv · Europe PMC · Crossref · checked 15 Sept 2026
Published18 Aug 2022bioRxivCited by 2 · OpenAlex ↗

An end-to-end workflow based on multimodal 3D imaging and machine learning for non-destructive diagnosis of grapevine trunk diseases

GrapevineField / plotMesh / voxelMRI / PETMultimodalX-ray / CTStem / branchTissueClassificationObject detection

Quantifying healthy and degraded inner tissues in plants is of great interest in agronomy, for example, to assess plant health and quality and monitor physiological traits or diseases. However, detecting functional and degraded plant tissues in-vivo without harming the plant is extremely challenging. New solutions are needed in ligneous and perennial species, for which the sustainability of plantations is crucial. To tackle this challenge, we developed a novel approach based on multimodal 3D imaging and Artificial Intelligence (AI)-based image processing that allowed a noninvasive diagnosis of inner tissues in living plants. The method was successfully applied to the grapevine (Vitis vinifera L.) in vineyards where sustainability was threatened by trunk diseases, while the sanitary status of vines cannot be ascertained without injuring the plants. By combining MRI and X-ray CT 3D imaging with an automatic voxel classification, we could discriminate intact, degraded, and white rot tissues with a mean global accuracy of over 91%. Each imaging modality contribution to tissue detection was evaluated, and we identified quantitative structural and physiological markers characterizing wood degradation steps. The combined study of inner tissue distribution versus external foliar symptom history demonstrated that white rot and intact tissue contents are key measurements in evaluating vines sanitary status. We finally proposed a model for an accurate trunk disease diagnosis in grapevine. This work opens new routes for precision agriculture and in-situ monitoring of wood quality and plant health across plant species.

Why it matches plant phenotyping methodsブドウ樹内部組織と病害状態を、MRI・X線CT・自動ボクセル分類によって非破壊的に定量する手法を開発・評価しており、植物表現型取得が研究の中心である。

abstractwe developed a novel approach based on multimodal 3D imaging and Artificial Intelligence (AI)-based image processing that allowed a noninvasive diagnosis of inner tissues in living plants
Reproduction assets foundThe paper's imaging datasets (MRI, X-ray CT, photographic volumes, annotations) are only available 'upon reasonable request', but the authors' extended Trainable Segmentation plugin used for the machine-learning voxel classification is explicitly open-source on GitHub.
Code · publicFernandez et al. 24 DATA AND CODE AVAILABILITY The datasets generated and analyzed during the current study are available from the corresponding author upon reasonable request. The extension of the Trainable Segmentation plugin is open-source, and available as a fork of Trainable Segmentation on GitHub: https://github.com/Rocsg/Trainable_Segmentation/tree/Hyperweka. . CC-BY-NC-ND 4.0 International license perpetuity. It is made available under a preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in The copyright holder for this this version posted February 3, 2023. ; https://doi.org/10.1101/2022.06.09.495457 doOpen asset ↗Rocsg/Trainable_Segmentation · Hyperwekapdf-raw-page:24 lines:1-16
Code / dataset availability confirmedEurope PMC · Crossref · checked 15 Sept 2026
Published1 Feb 2022Plant PhysiologyCited by 91 · OpenAlex ↗

X-ray microscopy enables multiscale high-resolution 3D imaging of plant cells, tissues, and organs

Laboratory / benchtopMicroscopyMultimodalX-ray / CTCell / cellular structureTissueWhole plant / canopy / plot / field2D/3D reconstructionSegmentationGrowth / development / phenology

Capturing complete internal anatomies of plant organs and tissues within their relevant morphological context remains a key challenge in plant science. While plant growth and development are inherently multiscale, conventional light, fluorescence, and electron microscopy platforms are typically limited to imaging of plant microstructure from small flat samples that lack a direct spatial context to, and represent only a small portion of, the relevant plant macrostructures. We demonstrate technical advances with a lab-based X-ray microscope (XRM) that bridge the imaging gap by providing multiscale high-resolution three-dimensional (3D) volumes of intact plant samples from the cell to the whole plant level. Serial imaging of a single sample is shown to provide sub-micron 3D volumes co-registered with lower magnification scans for explicit contextual reference. High-quality 3D volume data from our enhanced methods facilitate sophisticated and effective computational segmentation. Advances in sample preparation make multimodal correlative imaging workflows possible, where a single resin-embedded plant sample is scanned via XRM to generate a 3D cell-level map, and then used to identify and zoom in on sub-cellular regions of interest for high-resolution scanning electron microscopy. In total, we present the methodologies for use of XRM in the multiscale and multimodal analysis of 3D plant features using numerous economically and scientifically important plant systems.

Why it matches plant phenotyping methods植物の細胞から個体レベルの3D形態を取得するX線顕微鏡法と試料調製・計算セグメンテーションを中心に開発・提示しており、植物表現型取得手法が明確に主題である。

abstractWe demonstrate technical advances with a lab-based X-ray microscope (XRM) that bridge the imaging gap by providing multiscale high-resolution three-dimensional (3D) volumes of intact plant samples from the cell to the whole plant level.
Reproduction assets foundThe authors deposited fly-through animations of 2D image stacks and 3D volume rendering animations of the XRM scans shown in the paper's figures on figshare, directly reproducing this paper's plant phenotyping imaging data. No author analysis code or trained model checkpoints were explicitly deposited.
Dataset · publicCanada) was used for data integration, visualization, and animation of the scan data, and to export image data as 2D 16-bit Tag Image File Format (TIFF) stacks. Fly-through animations of 2D image stacks for scans shown in all Figures, as well as 3D volume rendering animations of selected scans, are available for download from ( https://figshare.com/s/944efc8832e47fd4f203 ). Image analysis and segmentation Data from XRM scans were segmented using Amira software and with the assistance of a Wacom tablet for manual segmentation, in addition to ORS Dragonfly Deep Learning Module 2021.1.0.977 which is free for noncommercial use. Segmentation for Figure 1D combined automated and manual methods in Open asset ↗figsharelines:87-114
Code / dataset availability confirmedEurope PMC · Crossref · bioRxiv · checked 15 Sept 2026
Published19 Dec 2020openRxivCited by 6 · OpenAlex ↗

X-ray microscopy enables multiscale high-resolution 3D imaging of plant cells, tissues, and organs

Laboratory / benchtopMicroscopyMultimodalX-ray / CTCell / cellular structureTissueWhole plant / canopy / plot / field2D/3D reconstructionSegmentationGrowth / development / phenology

Capturing complete internal anatomies of plant organs and tissues within their relevant morphological context remains a key challenge in plant science. While plant growth and development are inherently multiscale, conventional light, fluorescence, and electron microscopy platforms are typically limited to imaging of plant microstructure from small flat samples that lack direct spatial context to, and represent only a small portion of, the relevant plant macrostructures. We demonstrate technical advances with a lab-based X-ray microscope (XRM) that bridge the imaging gap by providing multiscale high-resolution 3D volumes of intact plant samples from the cell to whole plant level. Serial imaging of a single sample is shown to provide sub-micron 3D volumes co-registered with lower magnification scans for explicit contextual reference. High quality 3D volume data from our enhanced methods facilitate more sophisticated and effective computational segmentation and analyses than have previously been employed for X-ray based imaging. Advances in sample preparation make multimodal correlative imaging workflows possible, where a single resin-embedded plant sample is scanned via XRM to generate a 3D cell-level map, and then used to identify and zoom in on sub-cellular regions of interest for high resolution scanning electron microscopy. In total, we present the methodologies for use of XRM in the multiscale and multimodal analysis of 3D plant features using numerous economically and scientifically important plant systems.

Why it matches plant phenotyping methods植物試料の細胞から個体までを対象に、X線顕微鏡によるマルチスケール3D画像取得、試料調製、計算セグメンテーション、相関イメージングの方法論を中心に提示しており、植物形態の取得・解析法が明確に中心です。

abstractWe demonstrate technical advances with a lab-based X-ray microscope (XRM) that bridge the imaging gap by providing multiscale high-resolution 3D volumes of intact plant samples from the cell to whole plant level.
Reproduction assets foundThe preprint points to a public figshare collection containing the paper's high-resolution XRM image stacks ('flythroughs') and videos of the 3D plant datasets, which directly reproduce the paper's phenotyping imaging measurements. No author analysis code or trained model checkpoint is explicitly deposited; the deep-se
Dataset · publicof these improved techniques will 112 make a significant contribution to plant biology, expanding the reach of XRM as a 113 routine tool for 3D imaging for plant scientists. 114 115 116 RESULTS1 117 118 Meristem Biology 119 1 high-resolution image stacks (“flythroughs”) and videos portraying the 3D data sets can be found here: https://figshare.com/s/944efc8832e47fd4f203 . CC-BY-NC 4.0 International license available under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint this version posted December 22, 2020. ; https://doi.org/10.1101/2020.12.18.423480 doiOpen asset ↗figsharepdf-raw-page:4 lines:1-64
Code / dataset availability confirmedEurope PMC · Crossref · OpenAlex · checked 9 Sept 2026
Published30 Sept 2019PLOS ONECited by 6 · OpenAlex ↗

Comparison of feature point detectors for multimodal image registration in plant phenotyping

Chlorophyll fluorescenceMultimodalRGB / grayscaleLeafObject detectionImage / point-cloud registrationSegmentation

With the introduction of multi-camera systems in modern plant phenotyping new opportunities for combined multimodal image analysis emerge. Visible light (VIS), fluorescence (FLU) and near-infrared images enable scientists to study different plant traits based on optical appearance, biochemical composition and nutrition status. A straightforward analysis of high-throughput image data is hampered by a number of natural and technical factors including large variability of plant appearance, inhomogeneous illumination, shadows and reflections in the background regions. Consequently, automated segmentation of plant images represents a big challenge and often requires an extensive human-machine interaction. Combined analysis of different image modalities may enable automatisation of plant segmentation in "difficult" image modalities such as VIS images by utilising the results of segmentation of image modalities that exhibit higher contrast between plant and background, i.e. FLU images. For efficient segmentation and detection of diverse plant structures (i.e. leaf tips, flowers), image registration techniques based on feature point (FP) matching are of particular interest. However, finding reliable feature points and point pairs for differently structured plant species in multimodal images can be challenging. To address this task in a general manner, different feature point detectors should be considered. Here, a comparison of seven different feature point detectors for automated registration of VIS and FLU plant images is performed. Our experimental results show that straightforward image registration using FP detectors is prone to errors due to too large structural difference between FLU and VIS modalities. We show that structural image enhancement such as background filtering and edge image transformation significantly improves performance of FP algorithms. To overcome the limitations of single FP detectors, combination of different FP methods is suggested. We demonstrate application of our enhanced FP approach for automated registration of a large amount of FLU/VIS images of developing plant species acquired from high-throughput phenotyping experiments.

Why it matches plant phenotyping methods植物フェノタイピング用のVIS/FLU画像登録について、特徴点検出器を比較し、前処理と組合せ手法を評価する方法開発・検証研究であり、手法が中心的です。

abstractHere, a comparison of seven different feature point detectors for automated registration of VIS and FLU plant images is performed.
Reproduction assets foundThe authors publicly release example multimodal FLU/VIS plant images (original and manually segmented) together with a pre-compiled GUI demo tool implementing their FP registration analysis, via a dedicated IPK project page and a GitHub repository.
Code · publicA precompilied GUI tool demonstrating the peformance of different FP algorithms can be downloaded along with examples of multimodal plant images from https://github.com/ba-ipk/fpRegOpen asset ↗ba-ipk/fpReglines:261-304
Dataset · publicExamples of original (unfiltered) and manually segmented plant images along with the demo software are available from our project/paper dedicated page: http://ag-ba.ipk-gatersleben.de/fpreg.htmlOpen asset ↗lines:145-169
Code / dataset availability confirmedEurope PMC · OpenAlex · checked 10 Sept 2026
Published29 Apr 2019Plant MethodsCited by 10 · OpenAlex ↗

Comparison and extension of three methods for automated registration of multimodal plant images

Chlorophyll fluorescenceMultimodalRGB / grayscaleImage / point-cloud registration

With the introduction of high-throughput multisensory imaging platforms, the automatization of multimodal image analysis has become the focus of quantitative plant research. Due to a number of natural and technical reasons (e.g., inhomogeneous scene illumination, shadows, and reflections), unsupervised identification of relevant plant structures (i.e., image segmentation) represents a nontrivial task that often requires extensive human-machine interaction. Registration of multimodal plant images enables the automatized segmentation of 'difficult' image modalities such as visible light or near-infrared images using the segmentation results of image modalities that exhibit higher contrast between plant and background regions (such as fluorescent images). Furthermore, registration of different image modalities is essential for assessment of a consistent multiparametric plant phenotype, where, for example, chlorophyll and water content as well as disease- and/or stress-related pigmentation can simultaneously be studied at a local scale. To automatically register thousands of images, efficient algorithmic solutions for the unsupervised alignment of two structurally similar but, in general, nonidentical images are required. For establishment of image correspondences, different algorithmic approaches based on different image features have been proposed. The particularity of plant image analysis consists, however, of a large variability of shapes and colors of different plants measured at different developmental stages from different views. While adult plant shoots typically have a unique structure, young shoots may have a nonspecific shape that can often be hardly distinguished from the background structures. Consequently, it is not clear a priori what image features and registration techniques are suitable for the alignment of various multimodal plant images. Furthermore, dynamically measured plants may exhibit nonuniform movements that require application of nonrigid registration techniques. Here, we investigate three common techniques for registration of visible light and fluorescence images that rely on finding correspondences between (i) feature-points, (ii) frequency domain features, and (iii) image intensity information. The performance of registration methods is validated in terms of robustness and accuracy measured by a direct comparison with manually segmented images of different plants. Our experimental results show that all three techniques are sensitive to structural image distortions and require additional preprocessing steps including structural enhancement and characteristic scale selection. To overcome the limitations of conventional approaches, we develop an iterative algorithmic scheme, which allows it to perform both rigid and slightly nonrigid registration of high-throughput plant images in a fully automated manner.

Why it matches plant phenotyping methods植物のマルチモーダル画像登録アルゴリズムを比較・検証し、高スループット画像から一貫した表現型解析を可能にする手法を開発しているため。

abstractHere, we investigate three common techniques for registration of visible light and fluorescence images
Reproduction assets foundThe authors provide a public GUI software tool (mPIR) implementing the paper's FP/PC/INT multimodal plant image registration algorithms, together with example FLU/VIS/NIR plant images from the study, downloadable from their homepage.
Code · publica GUI software tool with examples of plant images is provided for direct download from our homepage; Footnote 1 a screen shot is shown in Fig. 10Open asset ↗lines:138-172
Dataset · publicExamples of FLU, VIS and NIR plant images are included in our online file repository.Open asset ↗lines:138-172
Code / dataset availability confirmedOpenAlex · Europe PMC · checked 10 Sept 2026
Published6 Feb 2019Plant MethodsCited by 54 · OpenAlex ↗

A spatio temporal spectral framework for plant stress phenotyping

Field / plotMultimodalRGB / grayscaleMultispectral / hyperspectralStereoWhole plant / canopy / plot / fieldClassification2D/3D reconstructionStress / disease detectionBiomass / plant weight

Recent advances in high throughput phenotyping have made it possible to collect large datasets following plant growth and development over time, and those in machine learning have made inferring phenotypic plant traits from such datasets possible. However, there remains a dirth of datasets following plant growth under stress conditions along with methods for inferring them using only remotely sensed data, especially under a combination of multiple stress factors such as drought, weeds and nutrient deficiency. Such stress factors and their combinations are commonly encountered during crop production and being able to accurately detect and treat such stress conditions in an automated and timely manner can provide a major boost to farm yields with minimal resource input. We present a generic framework for remote plant stress phenotyping that consists of a dataset with spatio-temporal-spectral data following sugarbeet crop growth under optimal, drought, low and surplus nitrogen fertilization, and weed stress conditions, along with a machine learning based methodology for systematically inferring these stress conditions from the remotely measured data. The dataset contains biweekly color images, infra-red stereo image pairs and hyperspectral camera images along with applied treatment parameters and environmental factors like temperature and humidity, collected over two months. We present a plant agnostic methodology for deriving plant trait indicators such as canopy cover, height, hyperspectral reflectance and vegetation indices along with a spectral 3D reconstruction of the plants from the raw data to serve as a benchmark. Additionally, we provide fresh and dry weight measurements for both the above (canopy) and below (beet) ground biomass at the end of the growing period to serve as indicators of expected yield. We further describe a data driven, machine learning based method to infer water, Nitrogen and weed stress using the derived plant trait indicators. We use the plant trait indicators to evaluate 8 different classification approaches from which the best classifier achieved a mean cross validation accuracy of $$\approx$$ 93, 76 and 83% for drought, nitrogen and weed stress severity classification respectively. We also show that our multi-modal approach significantly improves classifier performance over using any single modality. The presented framework and dataset can serve as a valuable reference for creating and comparing processing pipelines which extract plant trait indicators and infer prevalent stress factors from remote sensing data under a variety of environments and cropping conditions. These techniques can then be deployed on farm machinery or robots enabling automated, precise and timely corrective interventions for maximising yield.

Why it matches plant phenotyping methods植物ストレス表現型を推定するデータセット、マルチモーダル画像・分光計測、形質抽出、機械学習推定を一体化した汎用フレームワークであり、表現型取得・解析手法が研究の中心である。

abstractWe present a generic framework for remote plant stress phenotyping that consists of a dataset with spatio-temporal-spectral data
Reproduction assets foundThe paper releases its own plant stress phenotyping dataset (RGB, stereo IR, hyperspectral imagery, reference measurements) and accompanying pre-processing/classification software, both publicly available at author-provided URLs.
Dataset · publicThe images and reference data that support the findings of this study are available from ETH Zürich ASL Datasets Repository, “ https://projects.asl.ethz.ch/datasets/doku.php?id=2018plantstressphenotyping ”.Open asset ↗ETH Zürich ASL Datasets Repository · 2018plantstressphenotypinglines:367-481