Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Zhongang, Wang, Yubo, Sun, Qingping, Wang, Ruisi, Gu, Chenyang, Yin, Wanqi, Lin, Zhiqian, Yang, Zhitao, Wei, Chen, Qian, Oscar, Pang, Hui En, Shi, Xuanke, Deng, Kewang, Han, Xiaoyang, Chen, Zukai, Li, Jiaqi, Fan, Xiangyu, Deng, Hanming, Lu, Lewei, Li, Bo, Liu, Ziwei, Wang, Quan, Lin, Dahua, Yang, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Spatial Intelligence with Multimodal Foundation Models
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
von: Shi, Xuanke, et al.
Veröffentlicht: (2025)
von: Shi, Xuanke, et al.
Veröffentlicht: (2025)
WHAC: World-grounded Humans and Cameras
von: Yin, Wanqi, et al.
Veröffentlicht: (2024)
von: Yin, Wanqi, et al.
Veröffentlicht: (2024)
Demystifying Video Reasoning
von: Wang, Ruisi, et al.
Veröffentlicht: (2026)
von: Wang, Ruisi, et al.
Veröffentlicht: (2026)
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
von: Lin, Jing, et al.
Veröffentlicht: (2025)
von: Lin, Jing, et al.
Veröffentlicht: (2025)
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation
von: Yin, Wanqi, et al.
Veröffentlicht: (2025)
von: Yin, Wanqi, et al.
Veröffentlicht: (2025)
AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation
von: Sun, Qingping, et al.
Veröffentlicht: (2024)
von: Sun, Qingping, et al.
Veröffentlicht: (2024)
SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
von: Cai, Zhongang, et al.
Veröffentlicht: (2023)
von: Cai, Zhongang, et al.
Veröffentlicht: (2023)
From Pixels to Words -- Towards Native One-Vision Models at Scale
von: Diao, Haiwen, et al.
Veröffentlicht: (2026)
von: Diao, Haiwen, et al.
Veröffentlicht: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
von: Diao, Haiwen, et al.
Veröffentlicht: (2026)
von: Diao, Haiwen, et al.
Veröffentlicht: (2026)
Global Existence for General Systems of Isentropic Gas Dynamics via a Weighted Pressure Perturbation Approach
von: Chen, Kewang
Veröffentlicht: (2026)
von: Chen, Kewang
Veröffentlicht: (2026)
Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
von: Fan, Xiangyu, et al.
Veröffentlicht: (2025)
von: Fan, Xiangyu, et al.
Veröffentlicht: (2025)
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
von: Jiang, Jianping, et al.
Veröffentlicht: (2024)
von: Jiang, Jianping, et al.
Veröffentlicht: (2024)
ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
von: Shen, Wenhao, et al.
Veröffentlicht: (2025)
von: Shen, Wenhao, et al.
Veröffentlicht: (2025)
ACPO: Counteracting Likelihood Displacement in Vision-Language Alignment with Asymmetric Constraints
von: Huang, Kaili, et al.
Veröffentlicht: (2026)
von: Huang, Kaili, et al.
Veröffentlicht: (2026)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
von: Shen, Wenhao, et al.
Veröffentlicht: (2026)
von: Shen, Wenhao, et al.
Veröffentlicht: (2026)
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
von: Fan, Xiangyu, et al.
Veröffentlicht: (2024)
von: Fan, Xiangyu, et al.
Veröffentlicht: (2024)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
von: Chng, Yong Xien, et al.
Veröffentlicht: (2025)
von: Chng, Yong Xien, et al.
Veröffentlicht: (2025)
Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
von: Gu, Chenyang, et al.
Veröffentlicht: (2026)
von: Gu, Chenyang, et al.
Veröffentlicht: (2026)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving
von: Xu, Dongyang, et al.
Veröffentlicht: (2024)
von: Xu, Dongyang, et al.
Veröffentlicht: (2024)
Disco4D: Disentangled 4D Human Generation and Animation from a Single Image
von: Pang, Hui En, et al.
Veröffentlicht: (2024)
von: Pang, Hui En, et al.
Veröffentlicht: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
TF-DWGNet: A Directed Weighted Graph Neural Network with Tensor Fusion for Multi-Omics Cancer Subtype Classification
von: Yang, Tiantian, et al.
Veröffentlicht: (2025)
von: Yang, Tiantian, et al.
Veröffentlicht: (2025)
MOTGNN: Interpretable Graph Neural Networks for Multi-Omics Disease Classification
von: Yang, Tiantian, et al.
Veröffentlicht: (2025)
von: Yang, Tiantian, et al.
Veröffentlicht: (2025)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior
von: Lu, Junzhe, et al.
Veröffentlicht: (2025)
von: Lu, Junzhe, et al.
Veröffentlicht: (2025)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
von: Fan, Weichen, et al.
Veröffentlicht: (2025)
von: Fan, Weichen, et al.
Veröffentlicht: (2025)
Finiteness of non-decomposable critically 4 and 5-frustrated signed graphs
von: Wang, Zhiqian
Veröffentlicht: (2026)
von: Wang, Zhiqian
Veröffentlicht: (2026)
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
von: Zhang, Mengchen, et al.
Veröffentlicht: (2024)
von: Zhang, Mengchen, et al.
Veröffentlicht: (2024)
DynamiCtrl: Rethinking the Basic Structure and the Role of Text for High-quality Human Image Animation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
von: Lu, Hao, et al.
Veröffentlicht: (2025)
von: Lu, Hao, et al.
Veröffentlicht: (2025)
Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment
von: Siyao, Li, et al.
Veröffentlicht: (2024)
von: Siyao, Li, et al.
Veröffentlicht: (2024)
Study on the degradation of xylose residue in hydrothermal treatment process
von: Shuaifei Zhang, et al.
Veröffentlicht: (2026)
von: Shuaifei Zhang, et al.
Veröffentlicht: (2026)
Spectral Subspace Clustering for Attributed Graphs
von: Lin, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Lin, Xiaoyang, et al.
Veröffentlicht: (2024)
Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia
von: Zhou, Zhejian, et al.
Veröffentlicht: (2024)
von: Zhou, Zhejian, et al.
Veröffentlicht: (2024)
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scaling Spatial Intelligence with Multimodal Foundation Models
von: Cai, Zhongang, et al.
Veröffentlicht: (2025) -
ConsistCompose: Unified Multimodal Layout Control for Image Composition
von: Shi, Xuanke, et al.
Veröffentlicht: (2025) -
WHAC: World-grounded Humans and Cameras
von: Yin, Wanqi, et al.
Veröffentlicht: (2024) -
Demystifying Video Reasoning
von: Wang, Ruisi, et al.
Veröffentlicht: (2026) -
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)