AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Haocheng, Zheng, Juepeng, Yang, Zenghao, Du, Kaiqi, Xiao, Guilong, Pu, Gengmeng, Fu, Haohuan, Huang, Jianxi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913152044630016
author Li, Haocheng
Zheng, Juepeng
Yang, Zenghao
Du, Kaiqi
Xiao, Guilong
Pu, Gengmeng
Fu, Haohuan
Huang, Jianxi
author_facet Li, Haocheng
Zheng, Juepeng
Yang, Zenghao
Du, Kaiqi
Xiao, Guilong
Pu, Gengmeng
Fu, Haohuan
Huang, Jianxi
contents Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, and targeted harvesting. Reliable evaluation of agricultural visual grounding remains challenging because agricultural targets are often small, repetitive, occluded, or irregularly shaped, and instructions may refer to one, many, or no objects in an image. Evaluating this capability therefore requires jointly testing localization accuracy, target-set completeness, and existence-aware abstention. To address these challenges, we introduce \textbf{AgroVG}, a multi-source benchmark that formulates agricultural grounding as generalized set prediction: given an image and a referring expression, a model must return all matching target instances or abstain when no target is present. AgroVG contains 10{,}071 annotation-grounded image-query pairs from ten source datasets across six target families: crop/weed, fruit, wheat head, pest, plant disease, and tree canopy. It supports bounding-box grounding (T1) across all six families and instance-mask grounding (T2) on sources with reliable instance-level pixel annotations, with queries covering single-target, multi-target, and target-absent regimes. AgroVG further provides task-specific protocols for box-set matching and query-level mask coverage. Zero-shot evaluation of 26 model configurations spanning closed-source MLLMs, open-source VLMs, and specialized grounding systems reveals persistent gaps: the best multi-target Set-$F_1$ reaches only 0.35, and the best positive-query mask success rate at IoU@0.75 remains below 0.17. Data and code are available at https://anonymous.4open.science/r/AgroVG-5172/ .
format Preprint
id arxiv_https___arxiv_org_abs_2605_22034
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
Li, Haocheng
Zheng, Juepeng
Yang, Zenghao
Du, Kaiqi
Xiao, Guilong
Pu, Gengmeng
Fu, Haohuan
Huang, Jianxi
Computer Vision and Pattern Recognition
Artificial Intelligence
Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, and targeted harvesting. Reliable evaluation of agricultural visual grounding remains challenging because agricultural targets are often small, repetitive, occluded, or irregularly shaped, and instructions may refer to one, many, or no objects in an image. Evaluating this capability therefore requires jointly testing localization accuracy, target-set completeness, and existence-aware abstention. To address these challenges, we introduce \textbf{AgroVG}, a multi-source benchmark that formulates agricultural grounding as generalized set prediction: given an image and a referring expression, a model must return all matching target instances or abstain when no target is present. AgroVG contains 10{,}071 annotation-grounded image-query pairs from ten source datasets across six target families: crop/weed, fruit, wheat head, pest, plant disease, and tree canopy. It supports bounding-box grounding (T1) across all six families and instance-mask grounding (T2) on sources with reliable instance-level pixel annotations, with queries covering single-target, multi-target, and target-absent regimes. AgroVG further provides task-specific protocols for box-set matching and query-level mask coverage. Zero-shot evaluation of 26 model configurations spanning closed-source MLLMs, open-source VLMs, and specialized grounding systems reveals persistent gaps: the best multi-target Set-$F_1$ reaches only 0.35, and the best positive-query mask success rate at IoU@0.75 remains below 0.17. Data and code are available at https://anonymous.4open.science/r/AgroVG-5172/ .
title AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.22034