the exploration of technical and textual features of plant disease datasets for enhancement of machine learning and deep learning applications, as the case may be. 2. Methodology 2.1. Data Collection A set of five distinct plant disease datasets that are publicly available were acquired online via: • "plant_village dataset" - [https://github.com/spMohanty/PlantVillage-Dataset] • "a_database_of_leaf_images" - [https://data.mendeley.com/datasets/hb74ynkjcn/4] • "RoCoLe dataset" - [https://data.mendeley.com/datasets/c5yvn32dzg/2] • "FGVCx_cassava dataset" - [https://github.com/icassava/fgvcx-icassava/tree/master#fgvcx-cassava-disease-diagnosis] and • "paddy_doctor dataset" [https://ieee-datapor
Open resource ↗PlantVillage-Dataset · pdf-raw-page:4 lines:1-49Unverified paper record
Revealing GLCM Metric Variations across Plant Disease Dataset: A Comprehensive Examination and Future Prospects for Enhanced Deep Learning Applications
25 Apr 2024 · 10.20944/preprints202404.1566.v1
Abstract
The intricate relationship between Gray-Level Co-occurrence Matrix (GLCM) metrics and machine learning model performance underscores the need for rigorous dataset evaluation and selection protocols to ensure the reliability and generalizability of classification outcomes. This study involved a thorough examination of selected publicly available plant diseases datasets, with an emphasis on how well they performed as measured by GLCM metrics. After first classifying the datasets according to their GLCM metrics, dataset_2 (D2) and dataset_5 (D5) were found, respectively, to be the best-performing dataset in all GLCM analyses. The same datasets were then used to train deep learning models, and their classification performances were assessed. A noteworthy association was observed between the results of training deep learning models and the performance ratings derived from GLCM studies. More specifically, dataset_2 (D2) performed best in both GLCM analysis and deep learning model performance, indicating a strong correlation between the accuracy of classification and the textural qualities that GLCM captured. In the context of plant disease identification, in particular, these results highlight the significance of clearly defined dataset selection criteria in deep learning applications. Scholars can improve the accuracy and dependability of deep learning models for diagnosing plant diseases by giving preference to datasets with favorable GLCM metrics. The research also emphasizes the importance of texture features being taken into account in addition to conventional image features, highlighting the necessity of transparency and rigor in dataset selection procedures.
Plant phenotyping relevance
植物病害画像データセットのGLCM特徴量と分類性能を比較評価し、病害状態の画像ベース推定におけるデータセット選定基準を検討しているため、評価手法・ベンチマークが中心である。
abstractThis study involved a thorough examination of selected publicly available plant diseases datasets, with an emphasis on how well they performed as measured by GLCM metrics.
abstractThe research also emphasizes the importance of texture features being taken into account in addition to conventional image features, highlighting the necessity of transparency and rigor in dataset selection procedures.
Code and data availability
The paper's GLCM and deep learning analyses were performed on five publicly available plant disease image datasets, each cited in the methodology with a public URL. No author-generated code, models, or derived data are shared (the Data Availability Statement is boilerplate MDPI text with no deposit).
ancement of machine learning and deep learning applications, as the case may be. 2. Methodology 2.1. Data Collection A set of five distinct plant disease datasets that are publicly available were acquired online via: • "plant_village dataset" - [https://github.com/spMohanty/PlantVillage-Dataset] • "a_database_of_leaf_images" - [https://data.mendeley.com/datasets/hb74ynkjcn/4] • "RoCoLe dataset" - [https://data.mendeley.com/datasets/c5yvn32dzg/2] • "FGVCx_cassava dataset" - [https://github.com/icassava/fgvcx-icassava/tree/master#fgvcx-cassava-disease-diagnosis] and • "paddy_doctor dataset" [https://ieee-dataport.org/documents/paddy-doctor-visual-image-dataset-automated-paddy-disease-classific
Open resource ↗a_database_of_leaf_images · pdf-raw-page:4 lines:1-49e may be. 2. Methodology 2.1. Data Collection A set of five distinct plant disease datasets that are publicly available were acquired online via: • "plant_village dataset" - [https://github.com/spMohanty/PlantVillage-Dataset] • "a_database_of_leaf_images" - [https://data.mendeley.com/datasets/hb74ynkjcn/4] • "RoCoLe dataset" - [https://data.mendeley.com/datasets/c5yvn32dzg/2] • "FGVCx_cassava dataset" - [https://github.com/icassava/fgvcx-icassava/tree/master#fgvcx-cassava-disease-diagnosis] and • "paddy_doctor dataset" [https://ieee-dataport.org/documents/paddy-doctor-visual-image-dataset-automated-paddy-disease-classification-and-benchmarking] Preprints.org (www.preprints.org) | NOT PEER-RE
Open resource ↗RoCoLe dataset · pdf-raw-page:4 lines:1-49ease datasets that are publicly available were acquired online via: • "plant_village dataset" - [https://github.com/spMohanty/PlantVillage-Dataset] • "a_database_of_leaf_images" - [https://data.mendeley.com/datasets/hb74ynkjcn/4] • "RoCoLe dataset" - [https://data.mendeley.com/datasets/c5yvn32dzg/2] • "FGVCx_cassava dataset" - [https://github.com/icassava/fgvcx-icassava/tree/master#fgvcx-cassava-disease-diagnosis] and • "paddy_doctor dataset" [https://ieee-dataport.org/documents/paddy-doctor-visual-image-dataset-automated-paddy-disease-classification-and-benchmarking] Preprints.org (www.preprints.org) | NOT PEER-REVIEWED | Posted: 25 April 2024 doi:10.20944/preprints202404.1566.v1
Open resource ↗FGVCx_cassava dataset · pdf-raw-page:4 lines:1-49This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.