SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wufei, Ye, Luoxin, de Melo, Celso M, Chen, Jieneng, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
by: Guo, Weijie, et al.
Published: (2025)
by: Guo, Weijie, et al.
Published: (2025)
Thinking with Spatial Code for Physical-World Video Reasoning
by: Chen, Jieneng, et al.
Published: (2026)
by: Chen, Jieneng, et al.
Published: (2026)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)
by: Zhong, Shanshan, et al.
Published: (2025)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
by: Shi, Jiaxin, et al.
Published: (2026)
by: Shi, Jiaxin, et al.
Published: (2026)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
by: Sheung, Eddie Pokming, et al.
Published: (2025)
by: Sheung, Eddie Pokming, et al.
Published: (2025)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
by: Duan, Ruxiao, et al.
Published: (2023)
by: Duan, Ruxiao, et al.
Published: (2023)
GenEx: Generating an Explorable World
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
by: Ma, Wufei, et al.
Published: (2026)
by: Ma, Wufei, et al.
Published: (2026)
NOVUM: Neural Object Volumes for Robust Object Classification
by: Jesslen, Artur, et al.
Published: (2023)
by: Jesslen, Artur, et al.
Published: (2023)
Generative World Explorer
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
by: Guan, Yaohan, et al.
Published: (2026)
by: Guan, Yaohan, et al.
Published: (2026)
Pursuing Minimal Sufficiency in Spatial Reasoning
by: Guo, Yejie, et al.
Published: (2025)
by: Guo, Yejie, et al.
Published: (2025)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
by: Paul, Soumava, et al.
Published: (2026)
by: Paul, Soumava, et al.
Published: (2026)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
SITE: towards Spatial Intelligence Thorough Evaluation
by: Wang, Wenqi, et al.
Published: (2025)
by: Wang, Wenqi, et al.
Published: (2025)
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
by: Wang, Xiaoyan, et al.
Published: (2025)
by: Wang, Xiaoyan, et al.
Published: (2025)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Generating Images with 3D Annotations Using Diffusion Models
by: Ma, Wufei, et al.
Published: (2023)
by: Ma, Wufei, et al.
Published: (2023)
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
by: Gong, Ziyang, et al.
Published: (2025)
by: Gong, Ziyang, et al.
Published: (2025)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026)
by: Ma, Chuang, et al.
Published: (2026)
Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
by: He, Peiqi, et al.
Published: (2025)
by: He, Peiqi, et al.
Published: (2025)
How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
by: Dongfang, Zihao, et al.
Published: (2025)
by: Dongfang, Zihao, et al.
Published: (2025)
Similar Items
-
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025) -
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
by: Ma, Wufei, et al.
Published: (2024) -
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025) -
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024) -
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)