← Papers

Unverified paper record

A multi-agent vision-language debate framework for zero-shot crop disease diagnosis

Frontiers in Plant Science · 11 Aug 2026 · 10.3389/fpls.2026.1890016

Abstract

Accurate crop disease diagnosis is critical for agricultural productivity and food security, yet existing deep learning systems often struggle to generalize across visually similar diseases and varying environmental conditions. Recent Vision-Language Models (VLMs) have demonstrated promising zero-shot reasoning capabilities; however, most agricultural diagnostic systems still rely on isolated single-model predictions without collaborative reasoning or consensus mechanisms. In this work, we propose VIDA+PANDA, a multi-agent Vision-Language framework for zero-shot crop disease diagnosis. The framework consists of two stages: VIDA, where multiple VLM agents independently analyze crop leaf images to establish baseline performance, and the Peer-Anchored Named Deliberation Architecture (PANDA), which introduces a structured multi-round debate among a selected group of high-performing and architecturally diverse agents. During deliberation, agents exchange reasoning, critique peer predictions, and revise decisions through evidence-grounded discussion, while an anti-sycophancy mechanism discourages unsupported consensus shifts. Final predictions are generated through performance-weighted consensus voting. Experiments are conducted on the CDDM benchmark using seven heterogeneous VLMs from four independent providers, including two open-source models, under a fully zero-shot setting. A non-participant GPT-5 model serves as an independent judge to assess the final diagnostic predictions. Beyond conventional accuracy, the framework introduces three semantic measures: Semantic Label Similarity (SLS), which measures how semantically close a predicted crop-disease pair is to the ground truth and captures partial correctness overlooked by exact-match evaluation; Reasoning Specificity (RS), which measures how concretely an agent’s explanation references visual evidence such as lesion color, shape, texture, or margins; and Inter-Agent Reasoning Convergence (IRC), which measures the extent to which agents rely on similar visual evidence, capturing epistemic alignment independently of label correctness. Experimental results show that collaborative multiagent deliberation improves individual diagnostic performance and semantic alignment, with the largest gains observed among weaker participating agents. The findings also reveal important relationships between predictive accuracy, persuasive influence, and consensus formation in VLM-based agricultural diagnosis.

Plant phenotyping relevance

葉画像から植物病害状態を推定するマルチエージェント画像・言語フレームワークを開発し、ベンチマークで性能評価しており、病害表現型の取得・推定手法が中心である。

abstractwe propose VIDA+PANDA, a multi-agent Vision-Language framework for zero-shot crop disease diagnosis.
abstractagents exchange reasoning, critique peer predictions, and revise decisions through evidence-grounded discussion
abstractExperiments are conducted on the CDDM benchmark using seven heterogeneous VLMs from four independent providers
abstractReasoning Specificity (RS), which measures how concretely an agent’s explanation references visual evidence such as lesion color, shape, texture, or margins

Code and data availability

植物フェノタイピング解析を再現する公開資産であることを、入力本文と直接リンクから確認できなかったため保留しました。

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.