Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lundqvist, Lars, Ranario, Earl, Kamangir, Hamid, Yun, Heesup, Diepenbrock, Christine, Bailey, Brian N., Earles, J. Mason
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915932059729920
author Lundqvist, Lars
Ranario, Earl
Kamangir, Hamid
Yun, Heesup
Diepenbrock, Christine
Bailey, Brian N.
Earles, J. Mason
author_facet Lundqvist, Lars
Ranario, Earl
Kamangir, Hamid
Yun, Heesup
Diepenbrock, Christine
Bailey, Brian N.
Earles, J. Mason
contents Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex agricultural scenes remains highly sensitive to text prompt construction. We present a systematic prompt optimization framework evaluating four open-vocabulary detectors -- YOLO World, SAM3, Grounding DINO, and OWLv2 -- for cowpea flower and pod detection across synthetic and real field imagery. We decompose prompts into eight axes and conduct one-factor-at-a-time analysis followed by combinatorial optimization, revealing that models respond divergently to prompt structure: conditions that optimize one architecture can collapse another. Applying model-specific combinatorial prompts yields substantial gains over a naive species-name baseline, including +0.357 mAP@0.5 for YOLO World and +0.362 mAP@0.5 for OWLv2 on synthetic cowpea flower data. To evaluate cross-task generalization, we use an LLM to translate the discovered axis structure to a morphologically distinct target -- cowpea pods -- and compare against prompting using the discovered optimal structures from synthetic flower data. Crucially, prompt structures optimized exclusively on synthetic data transfer effectively to real-world fields: synthetic-pipeline prompts match or exceed those discovered on labeled real data for the majority of model-object combinations (flower: 0.374 vs. 0.353 for YOLO World; pod: 0.429 vs. 0.371 for SAM3). Our findings demonstrate that prompt engineering can substantially close the gap between zero-shot VFMs and supervised detectors without requiring manual annotation, and that optimal prompts are model-specific, non-obvious, and transferable across domains.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09920
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection
Lundqvist, Lars
Ranario, Earl
Kamangir, Hamid
Yun, Heesup
Diepenbrock, Christine
Bailey, Brian N.
Earles, J. Mason
Computer Vision and Pattern Recognition
Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex agricultural scenes remains highly sensitive to text prompt construction. We present a systematic prompt optimization framework evaluating four open-vocabulary detectors -- YOLO World, SAM3, Grounding DINO, and OWLv2 -- for cowpea flower and pod detection across synthetic and real field imagery. We decompose prompts into eight axes and conduct one-factor-at-a-time analysis followed by combinatorial optimization, revealing that models respond divergently to prompt structure: conditions that optimize one architecture can collapse another. Applying model-specific combinatorial prompts yields substantial gains over a naive species-name baseline, including +0.357 mAP@0.5 for YOLO World and +0.362 mAP@0.5 for OWLv2 on synthetic cowpea flower data. To evaluate cross-task generalization, we use an LLM to translate the discovered axis structure to a morphologically distinct target -- cowpea pods -- and compare against prompting using the discovered optimal structures from synthetic flower data. Crucially, prompt structures optimized exclusively on synthetic data transfer effectively to real-world fields: synthetic-pipeline prompts match or exceed those discovered on labeled real data for the majority of model-object combinations (flower: 0.374 vs. 0.353 for YOLO World; pod: 0.429 vs. 0.371 for SAM3). Our findings demonstrate that prompt engineering can substantially close the gap between zero-shot VFMs and supervised detectors without requiring manual annotation, and that optimal prompts are model-specific, non-obvious, and transferable across domains.
title Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.09920