Unverified paper record
From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
arXiv (Cornell University) · 10 Apr 2026 · 10.48550/arxiv.2604.09907
Abstract
To improve crop genetics, high-throughput, effective and comprehensive phenotyping is a critical prerequisite. While such tasks were traditionally performed manually, recent advances in multimodal foundation models, especially in vision-language models (VLMs), have enabled more automated and robust phenotypic analysis. However, plant science remains a particularly challenging domain for foundation models because it requires domain-specific knowledge, fine-grained visual interpretation, and complex biological and agronomic reasoning. To address this gap, we develop PlantXpert, an evidence-grounded multimodal reasoning benchmark for soybean and cotton phenotyping. Our benchmark provides a structured and reproducible framework for agronomic adaptation of VLMs, and enables controlled comparison between base models and their domain-adapted counterparts. We constructed a dataset comprising 385 digital images and more than 3,000 benchmark samples spanning key plant science domains including disease, pest control, weed management, and yield. The benchmark can assess diverse capabilities including visual expertise, quantitative reasoning, and multi-step agronomic reasoning. A total of 11 state-of-the-art VLMs were evaluated. The results indicate that task-specific fine-tuning leads to substantial improvement in accuracy, with models such as Qwen3-VL-4B and Qwen3-VL-30B achieving up to 78%. At the same time, gains from model scaling diminish beyond a certain capacity, generalization across soybean and cotton remains uneven, and quantitative as well as biologically grounded reasoning continue to pose substantial challenges. These findings suggest that PlantXpert can serve as a foundation for assessing evidence-grounded agronomic reasoning and for advancing multimodal model development in plant science.
Plant phenotyping relevance
PlantXpertは作物フェノタイピング向けの画像ベンチマークとVLM評価基盤を構築しており、表現型解析手法・データセットの開発が研究の中心である。
abstractwe develop PlantXpert, an evidence-grounded multimodal reasoning benchmark for soybean and cotton phenotyping.
abstractOur benchmark provides a structured and reproducible framework for agronomic adaptation of VLMs, and enables controlled comparison between base models and their domain-adapted counterparts.
abstractWe constructed a dataset comprising 385 digital images and more than 3,000 benchmark samples
Code and data availability
The supplied blocks describe the PlantXpert benchmark (385 images, 3,000+ QA pairs, Silver/Gold splits) but contain no public deposit, availability statement, or authors' URL for the dataset, images, annotations, fine-tuned model checkpoints, or analysis code. All URLs in the text are citations to prior work or generic
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.