TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Ming, Zhong, Jike, Zhao, Shitian, Zhang, Haoquan, Lin, Shaoheng, Lai, Yuxiang, Wei, Chen, Psounis, Konstantinos, Zhang, Kaipeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915601440571392
author Li, Ming
Zhong, Jike
Zhao, Shitian
Zhang, Haoquan
Lin, Shaoheng
Lai, Yuxiang
Wei, Chen
Psounis, Konstantinos
Zhang, Kaipeng
author_facet Li, Ming
Zhong, Jike
Zhao, Shitian
Zhang, Haoquan
Lin, Shaoheng
Lai, Yuxiang
Wei, Chen
Psounis, Konstantinos
Zhang, Kaipeng
contents The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet existing benchmarks fail to fully capture this advanced capability. Even Visual Search, the most common benchmark for current thinking-\textit{with}-images methods, tests only basic operations such as localization and cropping, offering little insight into more complex, dynamic, and tool-dependent reasoning. We introduce \textbf{TIR-Bench}, a comprehensive benchmark for evaluating agentic thinking-with-images across 13 diverse tasks, each requiring novel tool use for image processing and manipulation in chain-of-thought. We evaluate 22 multimodal large language models (MLLMs), from leading open-sourced and proprietary models to those with explicit tool-use augmentation. Results show that TIR-Bench is universally challenging, and strong performance requires genuine thinking-with-images capabilities. Finally, we present a pilot study comparing direct versus agentic fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Li, Ming
Zhong, Jike
Zhao, Shitian
Zhang, Haoquan
Lin, Shaoheng
Lai, Yuxiang
Wei, Chen
Psounis, Konstantinos
Zhang, Kaipeng
Computer Vision and Pattern Recognition
The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet existing benchmarks fail to fully capture this advanced capability. Even Visual Search, the most common benchmark for current thinking-\textit{with}-images methods, tests only basic operations such as localization and cropping, offering little insight into more complex, dynamic, and tool-dependent reasoning. We introduce \textbf{TIR-Bench}, a comprehensive benchmark for evaluating agentic thinking-with-images across 13 diverse tasks, each requiring novel tool use for image processing and manipulation in chain-of-thought. We evaluate 22 multimodal large language models (MLLMs), from leading open-sourced and proprietary models to those with explicit tool-use augmentation. Results show that TIR-Bench is universally challenging, and strong performance requires genuine thinking-with-images capabilities. Finally, we present a pilot study comparing direct versus agentic fine-tuning.
title TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.01833