RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jun, Feng, Jie, Chen, Long, Wang, Junhui, Liu, Zhicheng, Jin, Depeng, Li, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
by: Liu, Xiyan, et al.
Published: (2025)
by: Liu, Xiyan, et al.
Published: (2025)
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025)
by: Qiang, Chenhui, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark
by: Long, Hongyu, et al.
Published: (2026)
by: Long, Hongyu, et al.
Published: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
by: Li, Xiaomin, et al.
Published: (2026)
by: Li, Xiaomin, et al.
Published: (2026)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
by: Feng, Jie, et al.
Published: (2025)
by: Feng, Jie, et al.
Published: (2025)
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
by: Wang, Yandi, et al.
Published: (2026)
by: Wang, Yandi, et al.
Published: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
by: Tu, Chongjun, et al.
Published: (2025)
by: Tu, Chongjun, et al.
Published: (2025)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025)
by: Parikh, Chirag, et al.
Published: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
FunBench: Benchmarking Fundus Reading Skills of MLLMs
by: Wei, Qijie, et al.
Published: (2025)
by: Wei, Qijie, et al.
Published: (2025)
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
by: Xiao, Yicheng, et al.
Published: (2026)
by: Xiao, Yicheng, et al.
Published: (2026)
A Large-Scale Remote Sensing Dataset and VLM-based Algorithm for Fine-Grained Road Hierarchy Classification
by: Han, Ting, et al.
Published: (2026)
by: Han, Ting, et al.
Published: (2026)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Understanding Real-World Traffic Safety through RoadSafe365 Benchmark
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation
by: Wang, Bingtao, et al.
Published: (2026)
by: Wang, Bingtao, et al.
Published: (2026)
Exploiting Minority Pseudo-Labels for Semi-Supervised Fine-grained Road Scene Understanding
by: Hong, Yuting, et al.
Published: (2024)
by: Hong, Yuting, et al.
Published: (2024)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
by: Liu, Shaoyu, et al.
Published: (2025)
by: Liu, Shaoyu, et al.
Published: (2025)
Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
by: Muturi, Tanner, et al.
Published: (2025)
by: Muturi, Tanner, et al.
Published: (2025)
FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
by: Ji, Fengxian, et al.
Published: (2025)
by: Ji, Fengxian, et al.
Published: (2025)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
by: Zeng, Tong, et al.
Published: (2025)
by: Zeng, Tong, et al.
Published: (2025)
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
by: Chen, Xin-Sheng, et al.
Published: (2026)
by: Chen, Xin-Sheng, et al.
Published: (2026)
Similar Items
-
RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
by: Liu, Xiyan, et al.
Published: (2025) -
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025) -
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025) -
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025) -
Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark
by: Long, Hongyu, et al.
Published: (2026)