UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Siqi, Cai, Xinyu, Mei, Jianbiao, Deng, Nianchen, Cai, Pinlong, Wen, Licheng, Shen, Yufan, Yang, Xuemeng, Shi, Botian, Liu, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988984717312
author Li, Siqi
Cai, Xinyu
Mei, Jianbiao
Deng, Nianchen
Cai, Pinlong
Wen, Licheng
Shen, Yufan
Yang, Xuemeng
Shi, Botian
Liu, Yong
author_facet Li, Siqi
Cai, Xinyu
Mei, Jianbiao
Deng, Nianchen
Cai, Pinlong
Wen, Licheng
Shen, Yufan
Yang, Xuemeng
Shi, Botian
Liu, Yong
contents Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks typically rely on medium-resolution data, offering limited visual complexity. To bridge this gap, we introduce Ultra-high-resolution Reasoning Benchmark (UR-Bench), a benchmark designed to evaluate the reasoning capabilities of MLLMs under extreme visual information. UR-Bench comprises two major categories, Humanistic Scenes and Natural Scenes, covering four subsets of ultra-high-resolution images with distinct spatial structures and data sources. Each subset contains images ranging from hundreds of megapixels to gigapixels, accompanied by questions organized into three levels, enabling evaluation of models' reasoning capabilities in ultra-high-resolution scenarios. We further propose an agent-based framework in which a language model performs reasoning by invoking external visual tools. In addition, we introduce Semantic Abstraction and Retrieval tools that enable more efficient processing of ultra-high-resolution images. We evaluate state-of-the-art models using both an end-to-end MLLMs and our agent-based framework, demonstrating the effectiveness of our framework.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
Li, Siqi
Cai, Xinyu
Mei, Jianbiao
Deng, Nianchen
Cai, Pinlong
Wen, Licheng
Shen, Yufan
Yang, Xuemeng
Shi, Botian
Liu, Yong
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks typically rely on medium-resolution data, offering limited visual complexity. To bridge this gap, we introduce Ultra-high-resolution Reasoning Benchmark (UR-Bench), a benchmark designed to evaluate the reasoning capabilities of MLLMs under extreme visual information. UR-Bench comprises two major categories, Humanistic Scenes and Natural Scenes, covering four subsets of ultra-high-resolution images with distinct spatial structures and data sources. Each subset contains images ranging from hundreds of megapixels to gigapixels, accompanied by questions organized into three levels, enabling evaluation of models' reasoning capabilities in ultra-high-resolution scenarios. We further propose an agent-based framework in which a language model performs reasoning by invoking external visual tools. In addition, we introduce Semantic Abstraction and Retrieval tools that enable more efficient processing of ultra-high-resolution images. We evaluate state-of-the-art models using both an end-to-end MLLMs and our agent-based framework, demonstrating the effectiveness of our framework.
title UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.08748