AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ye, Zi, Wen, Yibin, Fan, Xiaoya, Zhang, Xinyu, Wu, Jing, Zeng, Kun, Mai, Zurong, Zhang, Jiarui, Shi, Bohan, Zheng, Juepeng, Huang, Jianxi, Lu, Yutong, Fu, Haohuan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917519893200896
author Ye, Zi
Wen, Yibin
Fan, Xiaoya
Zhang, Xinyu
Wu, Jing
Zeng, Kun
Mai, Zurong
Zhang, Jiarui
Shi, Bohan
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
author_facet Ye, Zi
Wen, Yibin
Fan, Xiaoya
Zhang, Xinyu
Wu, Jing
Zeng, Kun
Mai, Zurong
Zhang, Jiarui
Shi, Bohan
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
contents Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-answer correctness and provide limited support for assessing whether models can use external tools to complete precision-sensitive workflows. In this paper, we introduce AgroTools, a benchmark for evaluating tool-augmented multimodal agents in agriculture. AgroTools contains 539 question-answer instances paired with 1,097 heterogeneous agricultural images, spanning five task families and an executable environment of 14 agricultural tools. Each query is annotated with structured tool-use traces, enabling a dual-view evaluation of both process-level execution quality and outcome-level task success. We benchmark 9 open-source and 4 closed-source multimodal large language models on AgroTools. Results show that current models remain far from reliable in agricultural tool-use settings, with clear bottlenecks in tool planning, argument generation, execution recovery, and final-answer synthesis. We hope AgroTools will support future research on multimodal agents for high-precision agricultural applications. The benchmark and evaluation are available at https://huggingface.co/datasets/AgroTools/AgroTools.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
Ye, Zi
Wen, Yibin
Fan, Xiaoya
Zhang, Xinyu
Wu, Jing
Zeng, Kun
Mai, Zurong
Zhang, Jiarui
Shi, Bohan
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
Computer Vision and Pattern Recognition
Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-answer correctness and provide limited support for assessing whether models can use external tools to complete precision-sensitive workflows. In this paper, we introduce AgroTools, a benchmark for evaluating tool-augmented multimodal agents in agriculture. AgroTools contains 539 question-answer instances paired with 1,097 heterogeneous agricultural images, spanning five task families and an executable environment of 14 agricultural tools. Each query is annotated with structured tool-use traces, enabling a dual-view evaluation of both process-level execution quality and outcome-level task success. We benchmark 9 open-source and 4 closed-source multimodal large language models on AgroTools. Results show that current models remain far from reliable in agricultural tool-use settings, with clear bottlenecks in tool planning, argument generation, execution recovery, and final-answer synthesis. We hope AgroTools will support future research on multimodal agents for high-precision agricultural applications. The benchmark and evaluation are available at https://huggingface.co/datasets/AgroTools/AgroTools.
title AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22366