MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Sihan, Xu, Runsen, Xie, Yiman, Yang, Sizhe, Li, Mo, Lin, Jingli, Zhu, Chenming, Chen, Xiaochen, Duan, Haodong, Yue, Xiangyu, Lin, Dahua, Wang, Tai, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
Grounded 3D-LLM with Referent Tokens
by: Chen, Yilun, et al.
Published: (2024)
by: Chen, Yilun, et al.
Published: (2024)
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
by: Chen, Jiahe, et al.
Published: (2024)
by: Chen, Jiahe, et al.
Published: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
Unified Human-Scene Interaction via Prompted Chain-of-Contacts
by: Xiao, Zeqi, et al.
Published: (2023)
by: Xiao, Zeqi, et al.
Published: (2023)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
by: Ben, Qingwei, et al.
Published: (2025)
by: Ben, Qingwei, et al.
Published: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
by: Wang, Sizhe, et al.
Published: (2026)
by: Wang, Sizhe, et al.
Published: (2026)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
by: Zhuo, Jingming, et al.
Published: (2024)
by: Zhuo, Jingming, et al.
Published: (2024)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
by: Liu, Yuhong, et al.
Published: (2025)
by: Liu, Yuhong, et al.
Published: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023)
by: Huang, Haifeng, et al.
Published: (2023)
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
by: Tian, Yang, et al.
Published: (2024)
by: Tian, Yang, et al.
Published: (2024)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
by: Lei, Fei, et al.
Published: (2025)
by: Lei, Fei, et al.
Published: (2025)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
by: Li, Mo, et al.
Published: (2024)
by: Li, Mo, et al.
Published: (2024)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
EAG-PT: Emission-Aware Gaussians and Path Tracing for Diffuse Indoor Scene Reconstruction and Editing
by: Yang, Xijie, et al.
Published: (2026)
by: Yang, Xijie, et al.
Published: (2026)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
by: Fang, Xinyu, et al.
Published: (2025)
by: Fang, Xinyu, et al.
Published: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
Similar Items
-
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
by: Lin, Jingli, et al.
Published: (2025) -
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025) -
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025) -
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024) -
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)