Data accessibility Repository name: Harvard Dataverse Data identification number: https://doi.org/10.7910/DVN/FJ1DM1 Direct URL to data: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FJ1DM1
Open resource ↗Harvard Dataverse · doi:10.7910/DVN/FJ1DM1 · html-lines:1-91Unverified paper record
A comprehensive image dataset of jute diseases.
Data in brief · 28 Nov 2025 · 10.1016/j.dib.2025.112334
Abstract
This Data Descriptor presents the Jute Diseases Image Dataset; a curated collection of 1390 high-resolution images aimed at supporting the development of machine learning models for timely identification and accurate diagnosis of jute (Corchorus) plant diseases. The dataset is categorized into five classes: Dieback (300), Holed (300), Mosaic (240), Stem Soft Rot (270), and Fresh (280) representing healthy leaves. Images were captured under varied natural lighting and directional conditions across diverse jute cultivation areas to enhance model generalizability. A rigorous pre-processing pipeline was applied, including uniform resizing to 1024 × 1024 pixels and removal of duplicate images to ensure data integrity. The dataset is organized into two components: a raw, pre-processed set and an augmented train-test split version, enabling immediate use in machine learning workflows. Additionally, Grad-CAM and Guided Grad-CAM techniques were applied to sample images to visualize and validate model attention on disease-relevant regions. This resource addresses the lack of labelled jute disease imagery and supports timely disease management, particularly for stakeholders in Bangladesh and other major jute-producing regions.
Plant phenotyping relevance
植物病害症状を画像として収集・ラベル化したデータセットであり、病害状態の画像ベース表現型判定を支援することが中心です。
abstractThis Data Descriptor presents the Jute Diseases Image Dataset; a curated collection of 1390 high-resolution images aimed at supporting the development of machine learning models for timely identification and accurate diagnosis of jute (Corchorus) plant diseases.
abstractGrad-CAM and Guided Grad-CAM techniques were applied to sample images to visualize and validate model attention on disease-relevant regions.
Code and data availability
The paper's own jute disease image dataset (1390 labeled images, raw and augmented train/test splits) is publicly deposited in Harvard Dataverse with an explicit DOI and direct URL, matching an allowed URL. No separate analysis code or trained model checkpoint is publicly released.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.