3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wufei, Chen, Haoyu, Zhang, Guofeng, Chou, Yu-Cheng, Chen, Jieneng, de Melo, Celso M, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
by: Guo, Weijie, et al.
Published: (2025)
by: Guo, Weijie, et al.
Published: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)
by: Zhong, Shanshan, et al.
Published: (2025)
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
by: Shi, Jiaxin, et al.
Published: (2026)
by: Shi, Jiaxin, et al.
Published: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
by: Chen, Jieneng, et al.
Published: (2026)
by: Chen, Jieneng, et al.
Published: (2026)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
NOVUM: Neural Object Volumes for Robust Object Classification
by: Jesslen, Artur, et al.
Published: (2023)
by: Jesslen, Artur, et al.
Published: (2023)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
by: Duan, Ruxiao, et al.
Published: (2023)
by: Duan, Ruxiao, et al.
Published: (2023)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
by: Ma, Wufei, et al.
Published: (2026)
by: Ma, Wufei, et al.
Published: (2026)
Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape
by: Xu, Jiacong, et al.
Published: (2023)
by: Xu, Jiacong, et al.
Published: (2023)
Generative World Explorer
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
by: Sheung, Eddie Pokming, et al.
Published: (2025)
by: Sheung, Eddie Pokming, et al.
Published: (2025)
Generating Images with 3D Annotations Using Diffusion Models
by: Ma, Wufei, et al.
Published: (2023)
by: Ma, Wufei, et al.
Published: (2023)
EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
Embracing Massive Medical Data
by: Chou, Yu-Cheng, et al.
Published: (2024)
by: Chou, Yu-Cheng, et al.
Published: (2024)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
by: Guan, Yaohan, et al.
Published: (2026)
by: Guan, Yaohan, et al.
Published: (2026)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
by: Paul, Soumava, et al.
Published: (2026)
by: Paul, Soumava, et al.
Published: (2026)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Captain Safari: A World Engine with Pose-Aligned 3D Memory
by: Chou, Yu-Cheng, et al.
Published: (2025)
by: Chou, Yu-Cheng, et al.
Published: (2025)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Pursuing Minimal Sufficiency in Spatial Reasoning
by: Guo, Yejie, et al.
Published: (2025)
by: Guo, Yejie, et al.
Published: (2025)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
by: Wang, Jiayu, et al.
Published: (2025)
by: Wang, Jiayu, et al.
Published: (2025)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
Name That Part: 3D Part Segmentation and Naming
by: Paul, Soumava, et al.
Published: (2025)
by: Paul, Soumava, et al.
Published: (2025)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
by: Zhang, Guofeng, et al.
Published: (2026)
by: Zhang, Guofeng, et al.
Published: (2026)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
by: Chen, Yixiong, et al.
Published: (2024)
by: Chen, Yixiong, et al.
Published: (2024)
GenEx: Generating an Explorable World
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
by: Chen, Yixiong, et al.
Published: (2025)
by: Chen, Yixiong, et al.
Published: (2025)
M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
by: Weng, Ju-Hsuan, et al.
Published: (2025)
by: Weng, Ju-Hsuan, et al.
Published: (2025)
Similar Items
-
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
by: Ma, Wufei, et al.
Published: (2025) -
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025) -
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025) -
DINeMo: Learning Neural Mesh Models with no 3D Annotations
by: Guo, Weijie, et al.
Published: (2025) -
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)