Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Qingmei, Zhang, Yang, Mai, Zurong, Chen, Yuhang, Lou, Shuohong, Huang, Henglian, Zhang, Jiarui, Zhang, Zhiwei, Wen, Yibin, Li, Weijia, Fu, Haohuan, Huang, Jianxi, Zheng, Juepeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909734814089216
author Li, Qingmei
Zhang, Yang
Mai, Zurong
Chen, Yuhang
Lou, Shuohong
Huang, Henglian
Zhang, Jiarui
Zhang, Zhiwei
Wen, Yibin
Li, Weijia
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
author_facet Li, Qingmei
Zhang, Yang
Mai, Zurong
Chen, Yuhang
Lou, Shuohong
Huang, Henglian
Zhang, Jiarui
Zhang, Zhiwei
Wen, Yibin
Li, Weijia
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
contents Large Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diversity in the dataset and oversimplified task design. To bridge this gap, we introduce AgroMind, a comprehensive agricultural remote sensing benchmark covering four task dimensions: spatial perception, object understanding, scene understanding, and scene reasoning, with a total of 13 task types, ranging from crop identification and health monitoring to environmental analysis. We curate a high-quality evaluation set by integrating eight public datasets and one private farmland plot dataset, containing 27,247 QA pairs and 19,615 images. The pipeline begins with multi-source data pre-processing, including collection, format standardization, and annotation refinement. We then generate a diverse set of agriculturally relevant questions through the systematic definition of tasks. Finally, we employ LMMs for inference, generating responses, and performing detailed examinations. We evaluated 20 open-source LMMs and 4 closed-source models on AgroMind. Experiments reveal significant performance gaps, particularly in spatial reasoning and fine-grained recognition, it is notable that human performance lags behind several leading LMMs. By establishing a standardized evaluation framework for agricultural RS, AgroMind reveals the limitations of LMMs in domain knowledge and highlights critical challenges for future work. Data and code can be accessed at https://rssysu.github.io/AgroMind/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
Li, Qingmei
Zhang, Yang
Mai, Zurong
Chen, Yuhang
Lou, Shuohong
Huang, Henglian
Zhang, Jiarui
Zhang, Zhiwei
Wen, Yibin
Li, Weijia
Fu, Haohuan
Huang, Jianxi
Zheng, Juepeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diversity in the dataset and oversimplified task design. To bridge this gap, we introduce AgroMind, a comprehensive agricultural remote sensing benchmark covering four task dimensions: spatial perception, object understanding, scene understanding, and scene reasoning, with a total of 13 task types, ranging from crop identification and health monitoring to environmental analysis. We curate a high-quality evaluation set by integrating eight public datasets and one private farmland plot dataset, containing 27,247 QA pairs and 19,615 images. The pipeline begins with multi-source data pre-processing, including collection, format standardization, and annotation refinement. We then generate a diverse set of agriculturally relevant questions through the systematic definition of tasks. Finally, we employ LMMs for inference, generating responses, and performing detailed examinations. We evaluated 20 open-source LMMs and 4 closed-source models on AgroMind. Experiments reveal significant performance gaps, particularly in spatial reasoning and fine-grained recognition, it is notable that human performance lags behind several leading LMMs. By establishing a standardized evaluation framework for agricultural RS, AgroMind reveals the limitations of LMMs in domain knowledge and highlights critical challenges for future work. Data and code can be accessed at https://rssysu.github.io/AgroMind/.
title Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.12207