← Papers

Unverified paper record

SAM-CLIP-Thermal: Leveraging large multimodal models for reliable and scalable annotation in thermal image segmentation for field plant phenotyping.

Plant Phenomics · 12 Aug 2026 · 10.1016/j.plaphe.2026.100264

Abstract

Thermal imaging enables non-invasive assessment of canopy temperature, an essential indicator of plant stress, yet the lack of color cues and strong shadow interference make plant segmentation in thermal images difficult. Recent advances in foundation models have demonstrated improved performance and generalizability across applications, showing promise for domain-specific applications with limited annotated datasets such as plant segmentation in thermal images. This study investigates large multimodal models (LMMs) for thermal image segmentation in plant phenotyping. Building upon the SAM-CLIP framework, we design a unified pipeline spanning zero-shot inference, few-shot and low-shot fine-tuning, and active learning to maximize accuracy with minimal supervision. Evaluations on two thermal datasets, LadyBird Brassica and UGA Brassica, demonstrate robust performance after minimal adaptation across both datasets and superior performance compared with baselines, achieving mIoU D values of 97.54% on the LadyBird dataset and 76.94 % on the UGA dataset. We also release the resulting thermal segmentation annotations to support community benchmarking and reproducible research, highlighting the potential of LMMs to enable scalable, high-quality dataset construction for field phenotyping. The released datasets can be found at: https://cornell.box.com/s/dh69xf84464yrc1vlws92l1tflx7qa89

Plant phenotyping relevance

熱画像から植物を分割する手法を開発・評価し、植物フェノタイピング用データセットとアノテーションも公開しているため、フェノタイピング手法が中心的である。

abstractThis study investigates large multimodal models (LMMs) for thermal image segmentation in plant phenotyping.
abstractwe design a unified pipeline spanning zero-shot inference, few-shot and low-shot fine-tuning, and active learning to maximize accuracy with minimal supervision.
abstractWe also release the resulting thermal segmentation annotations to support community benchmarking and reproducible research

Code and data availability

The authors publicly released the paper-specific thermal segmentation annotations (20,538 LadyBird masks and 37,790 UGA masks) via a Cornell Box link stated in the abstract, results, and data availability statement. No author analysis code or trained model checkpoints are explicitly released; the mmsegmentation GitHub/

Datasetpublic

we generated and publicly released segmentation annotations for the complete LadyBird and UGA thermal image datasets using the best-performing SAM-CLIP model. Specifically, the final model obtained through the multi-round training process was used to generate 20,538 masks for the LadyBird dataset and 37,790 masks for the UGA dataset. Details of the generated annotations are provided in Supplementary Fig. S1 , and both annotated datasets are publicly available at: https://cornell.box.com/s/dh69xf84464yrc1vlws92l1tflx7qa89

Open resource ↗lines:220-232

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.