Unverified paper record
SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
arXiv (Cornell University) · 29 Mar 2026 · 10.48550/arxiv.2603.27519
Abstract
Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, organs, growth stages, and field conditions. General-purpose vision foundation models offer a natural route to label efficiency, but their web-scale pretraining objectives transfer weakly to agricultural imagery, where semantics are often determined by fine organ geometry inside repetitive, texture-dominated scenes. We introduce SPROUT, a diffusion foundation model for multi-crop plant phenotyping. SPROUT learns from 2.6 million unlabeled open-field images (MCD-2.6M) using a pixel-space Diffusion Transformer, and selects transferable features with a label-free effective-rank criterion over denoising timesteps. This design shifts pretraining from crop-based invariance to structure-preserving denoising, making the representation better aligned with dense phenotyping tasks. We evaluate SPROUT across dense phenotyping tasks, including organ segmentation, crop-weed parsing, depth estimation, and counting. SPROUT consistently improves over strong web-pretrained baselines, with the largest gains on dense structural prediction, and shows favorable label and compute efficiency compared with general-purpose and crop-specific foundation models. The source code and MCD-2.6M dataset are publicly available.
Plant phenotyping relevance
植物フェノタイピング向けの拡散基盤モデルを開発し、複数の作物・器官に対する画像ベースの構造推定タスクで評価しているため、表現学習法とデータセットが中心的な方法論的貢献である。
abstractWe introduce SPROUT, a diffusion foundation model for multi-crop plant phenotyping.
abstractWe evaluate SPROUT across dense phenotyping tasks, including organ segmentation, crop-weed parsing, depth estimation, and counting.
abstractThe source code and MCD-2.6M dataset are publicly available.
Code and data availability
The paper states that source code and the MCD-2.6M dataset are publicly available on GitHub and Hugging Face, but no concrete URLs are provided in the supplied blocks and the allowed_urls list is empty, so no paper-specific public asset can be verified or linked.
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.