MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jingli, Xu, Runsen, Zhu, Shaohao, Yang, Sihan, Cao, Peizhou, Ran, Yunlong, Hu, Miao, Zhu, Chenming, Xie, Yiman, Long, Yilin, Hu, Wenbo, Lin, Dahua, Wang, Tai, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
Grounded 3D-LLM with Referent Tokens
by: Chen, Yilun, et al.
Published: (2024)
by: Chen, Yilun, et al.
Published: (2024)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
by: Huang, Wensi, et al.
Published: (2025)
by: Huang, Wensi, et al.
Published: (2025)
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
by: Chen, Jiahe, et al.
Published: (2024)
by: Chen, Jiahe, et al.
Published: (2024)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
Conditional Neural Video Coding with Spatial-Temporal Super-Resolution
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
Unified Human-Scene Interaction via Prompted Chain-of-Contacts
by: Xiao, Zeqi, et al.
Published: (2023)
by: Xiao, Zeqi, et al.
Published: (2023)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
Learning H-Infinity Locomotion Control
by: Long, Junfeng, et al.
Published: (2024)
by: Long, Junfeng, et al.
Published: (2024)
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
by: Ben, Qingwei, et al.
Published: (2025)
by: Ben, Qingwei, et al.
Published: (2025)
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
by: Li, Bozheng, et al.
Published: (2025)
by: Li, Bozheng, et al.
Published: (2025)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025)
by: Peng, Liyang, et al.
Published: (2025)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
by: Liu, Xianjie, et al.
Published: (2026)
by: Liu, Xianjie, et al.
Published: (2026)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
by: Zhang, Yuanhan, et al.
Published: (2025)
by: Zhang, Yuanhan, et al.
Published: (2025)
IBVC: Interpolation-driven B-frame Video Compression
by: Xu, Chenming, et al.
Published: (2023)
by: Xu, Chenming, et al.
Published: (2023)
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
by: Tian, Yang, et al.
Published: (2024)
by: Tian, Yang, et al.
Published: (2024)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Similar Items
-
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
by: Yang, Sihan, et al.
Published: (2025) -
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025) -
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025) -
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025) -
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)