d 86.06%, respectively. Ablation experiments are conducted on CDwPK-VQA to evaluate the effectiveness of various modules, including coattention, MUTAN, and BiBa. These experiments demonstrate that ILCD exhibits the highest level of accuracy, performance, and value in the field of agriculture. The source codes can be accessed at https://github.com/SdustZYP/ILCD-master/tree/main. status released display-pdf yes is-olf no is-manuscript no is-preprint no is-journal-matter no is-scanned no is-retracted no Received 2024 May 24; Revised 2024 Oct 18; Accepted 2024 Nov 12; Collection date 2024. Introduction The Food and Agriculture Organization of the United Nations has reports that diseases are resp
Open resource ↗SdustZYP/ILCD-master · lines:1-26Unverified paper record
Informed-Learning-Guided Visual Question Answering Model of Crop Disease.
Plant phenomics (Washington, D.C.) · 16 Dec 2024 · 10.34133/plantphenomics.0277
Abstract
In contemporary agriculture, experts develop preventative and remedial strategies for various disease stages in diverse crops. Decision-making regarding the stages of disease occurrence exceeds the capabilities of single-image tasks, such as image classification and object detection. Consequently, research now focuses on training visual question answering (VQA) models. However, existing studies concentrate on identifying disease species rather than formulating questions that encompass crucial multiattributes. Additionally, model performance is susceptible to the model structure and dataset biases. To address these challenges, we construct the informed-learning-guided VQA model of crop disease (ILCD). ILCD improves model performance by integrating coattention, a multimodal fusion model (MUTAN), and a bias-balancing (BiBa) strategy. To facilitate the investigation of various visual attributes of crop diseases and the determination of disease occurrence stages, we construct a new VQA dataset called the Crop Disease Multi-attribute VQA with Prior Knowledge (CDwPK-VQA). This dataset contains comprehensive information on various visual attributes such as shape, size, status, and color. We expand the dataset by integrating prior knowledge into CDwPK-VQA to address performance challenges. Comparative experiments are conducted by ILCD on the VQA-v2, VQA-CP v2, and CDwPK-VQA datasets, achieving accuracies of 68.90%, 49.75%, and 86.06%, respectively. Ablation experiments are conducted on CDwPK-VQA to evaluate the effectiveness of various modules, including coattention, MUTAN, and BiBa. These experiments demonstrate that ILCD exhibits the highest level of accuracy, performance, and value in the field of agriculture. The source codes can be accessed at https://github.com/SdustZYP/ILCD-master/tree/main.
Plant phenotyping relevance
作物病害の視覚属性と発生段階を画像から推定するVQAモデルと専用データセットを開発しており、植物状態の表現型推定手法が研究の中心である。
abstractwe construct the informed-learning-guided VQA model of crop disease (ILCD).
abstractwe construct a new VQA dataset called the Crop Disease Multi-attribute VQA with Prior Knowledge (CDwPK-VQA).
abstractThis dataset contains comprehensive information on various visual attributes such as shape, size, status, and color.
Code and data availability
The paper's authors publicly release both the ILCD analysis code and the paper-specific CDwPK-VQA dataset (crop disease images with question–answer annotations) via GitHub URLs stated in the article.
gnment between the question text information and image region features. This process results prior knowledge dataset comprising 272 images and 2,180 questions. CDwPK-VQA integrates prior knowledge to expand the dataset and regulate the learning behavior of the model, as shown in Fig. 2 . The dataset of CDwPK-VQA is available at https://github.com/SdustZYP/CDwPK-VQA/tree/main. The ILCD model This research constructs a novel ILCD. The model architecture of ILCD is shown in Fig. 3 , and divided into the following steps: (a) Image features V and question features Q are extracted using a pretrained feature extraction model. (b) The coattention mechanism captures the interaction between the image
Open resource ↗SdustZYP/CDwPK-VQA · lines:52-91This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.