AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Yibin, Li, Qingmei, Ye, Zi, Zhang, Jiarui, Fan, Xiaoya, Mai, Zurong, Wu, Jing, Lou, Shuohong, Chen, Yuhang, Huang, Henglian, Zhang, Yang, Gu, Defeng, Zhao, Lingyuan, Lu, Yutong, Fu, Haohuan, Huang, Jianxi, Zheng, Juepeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910227378470912
author Wen, Yibin
Li, Qingmei
Ye, Zi
Zhang, Jiarui
Fan, Xiaoya
Mai, Zurong
Wu, Jing
Lou, Shuohong
Chen, Yuhang
Huang, Henglian
Zhang, Yang
Gu, Defeng
Zhao, Lingyuan
Lu, Yutong
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
author_facet Wen, Yibin
Li, Qingmei
Ye, Zi
Zhang, Jiarui
Fan, Xiaoya
Mai, Zurong
Wu, Jing
Lou, Shuohong
Chen, Yuhang
Huang, Henglian
Zhang, Yang
Gu, Defeng
Zhao, Lingyuan
Lu, Yutong
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
contents Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities hold great promise for applications such as precision farming, crop monitoring, pest detection, and environmental sustainability. However, while several Visual Question Answering (VQA) datasets and benchmarks have been developed to assess VLM performance, they often fail to effectively evaluate the critical reasoning and problem-solving skills needed in complex agricultural contexts. To address this gap, we introduce AgroCoT, a VQA dataset that integrates Chain-of-Thought (CoT) reasoning, specifically designed to evaluate the reasoning capabilities of VLMs. With 4,759 carefully curated samples, AgroCoT provides a comprehensive and robust evaluation of reasoning abilities, particularly in zero-shot scenarios, focusing on the models' ability to engage in logical reasoning and effective problem-solving. Our evaluation of 30 representative VLMs, including both proprietary and open-source models, reveals a gap in their reasoning capabilities, which underscores the importance of incorporating CoT for assessments. Our dataset is available at https://huggingface.co/datasets/AgroCoT/AgroCoT.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
Wen, Yibin
Li, Qingmei
Ye, Zi
Zhang, Jiarui
Fan, Xiaoya
Mai, Zurong
Wu, Jing
Lou, Shuohong
Chen, Yuhang
Huang, Henglian
Zhang, Yang
Gu, Defeng
Zhao, Lingyuan
Lu, Yutong
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
Artificial Intelligence
Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities hold great promise for applications such as precision farming, crop monitoring, pest detection, and environmental sustainability. However, while several Visual Question Answering (VQA) datasets and benchmarks have been developed to assess VLM performance, they often fail to effectively evaluate the critical reasoning and problem-solving skills needed in complex agricultural contexts. To address this gap, we introduce AgroCoT, a VQA dataset that integrates Chain-of-Thought (CoT) reasoning, specifically designed to evaluate the reasoning capabilities of VLMs. With 4,759 carefully curated samples, AgroCoT provides a comprehensive and robust evaluation of reasoning abilities, particularly in zero-shot scenarios, focusing on the models' ability to engage in logical reasoning and effective problem-solving. Our evaluation of 30 representative VLMs, including both proprietary and open-source models, reveals a gap in their reasoning capabilities, which underscores the importance of incorporating CoT for assessments. Our dataset is available at https://huggingface.co/datasets/AgroCoT/AgroCoT.
title AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
topic Artificial Intelligence
url https://arxiv.org/abs/2511.23253