← Papers

Unverified paper record

PlanText: Gradually Masked Guidance to Align Image Phenotypes with Trait Descriptions for Plant Disease Texts

Plant Phenomics · 26 Nov 2024 · 10.34133/plantphenomics.0272

Abstract

Plant diseases are a critical driver of the global food crisis. The integration of advanced artificial intelligence technologies can substantially enhance plant disease diagnostics. However, current methods for early and complex detection remain challenging. Employing multimodal technologies, akin to medical artificial intelligence diagnostics that combine diverse data types, may offer a more effective solution. Presently, the reliance on single-modal data predominates in plant disease research, which limits the scope for early and detailed diagnosis. Consequently, developing text modality generation techniques is essential for overcoming the limitations in plant disease recognition. To this end, we propose a method for aligning plant phenotypes with trait descriptions, which diagnoses text by progressively masking disease images. First, for training and validation, we annotate 5,728 disease phenotype images with expert diagnostic text and provide annotated text and trait labels for 210,000 disease images. Then, we propose a PhenoTrait text description model, which consists of global and heterogeneous feature encoders as well as switching-attention decoders, for accurate context-aware output. Next, to generate a more phenotypically appropriate description, we adopt 3 stages of embedding image features into semantic structures, which generate characterizations that preserve trait features. Finally, our experimental results show that our model outperforms several frontier models in multiple trait descriptions, including the larger models GPT-4 and GPT-4o. Our code and dataset are available at https://plantext.samlab.cn/.

Plant phenotyping relevance

植物病害画像から表現型・形質記述を生成するモデルと注釈付きデータセットを開発し、性能比較まで行っており、植物表現型の取得・抽出が研究の中心である。

abstractwe propose a method for aligning plant phenotypes with trait descriptions, which diagnoses text by progressively masking disease images.
abstractwe annotate 5,728 disease phenotype images with expert diagnostic text and provide annotated text and trait labels for 210,000 disease images.
abstractour experimental results show that our model outperforms several frontier models in multiple trait descriptions

Code and data availability

The paper's plant disease image–text dataset (17,183 annotated images, 103,098 labels, 5,728 expert texts) and the PhenoTrait/PlanText code are publicly released at the authors' site https://plantext.samlab.cn, and the authors' data annotation platform code is public on GitHub at https://github.com/kej-shas/data-angles

Codepublic

we develop a data annotation platform (the platform has built-in functions such as image display, text translation, and exporting of Word files; the code is at https://github.com/kej-shas/data-annotations )

Open resource ↗kej-shas/data-annotations · lines:77-85

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.