Unverified paper record
Vision–Language Models for Rice Pathology: From Diagnosis to Dialogue
24 Oct 2025 · 10.21203/rs.3.rs-7761035/v1
Abstract
Abstract Rice production is severely affected by major diseases such as bacterial panicle blight, bacterial leaf blight, and leaf blast, whose overlapping symptoms make diagnosis difficult in the field. While expert assessment remains the gold standard, access is limited for many farmers. Recent advances in vision–language models (VLMs) provide new opportunities for automated diagnosis and farmer-oriented dialogue, but their performance and safety in agricultural settings require careful evaluation. In this study, three VLMs (GPT-4o, Gemma3:27b, and Qwen2.5VL:72b) were evaluated. Performance was measured using recall, F1-score, and accuracy, with significance tested by McNemar’s test and stability assessed via 5,000 bootstrap resampling iterations. GPT-4o achieved the highest global accuracy (56.0%), followed by Gemma3:27b (44.7%), while Qwen2.5VL:72b lagged substantially (14.7%). At the disease level, bacterial leaf blight was consistently identified with high accuracy, while bacterial panicle blight and leaf blast were more difficult. Bootstrap analyses confirmed the robustness of these differences, showing a stable advantage for GPT-4o over Gemma3:27b in leaf blast, a narrower margin in bacterial leaf blight, and negligible difference in panicle blight. Beyond classification, the advisory dialogue module generated case-specific, clear, and consistently safe recommendations. These findings highlight the potential of VLMs, when guided by domain-specific prompting and a safety-first dialogue layer, to support both automated diagnosis and farmer-facing decision support in rice pathology.
Plant phenotyping relevance
イネ病害の症状を対象に、視覚言語モデルによる自動診断性能を比較・検証しており、植物の病害状態を直接推定する計算的方法が研究の中心である。
abstractRecent advances in vision–language models (VLMs) provide new opportunities for automated diagnosis and farmer-oriented dialogue, but their performance and safety in agricultural settings require careful evaluation.
abstractIn this study, three VLMs (GPT-4o, Gemma3:27b, and Qwen2.5VL:72b) were evaluated.
abstractPerformance was measured using recall, F1-score, and accuracy, with significance tested by McNemar’s test and stability assessed via 5,000 bootstrap resampling iterations.
Code and data availability
植物フェノタイピング解析を再現する公開資産であることを、入力本文と直接リンクから確認できなかったため保留しました。
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.