Space3D-Bench: Spatial 3D Question Answering Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Szymanska, Emilia, Dusmanu, Mihai, Buurlage, Jan-Willem, Rad, Mahdi, Pollefeys, Marc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
von: Qu, Kevin, et al.
Veröffentlicht: (2026)
von: Qu, Kevin, et al.
Veröffentlicht: (2026)
MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References
von: Bösiger, Lukas, et al.
Veröffentlicht: (2024)
von: Bösiger, Lukas, et al.
Veröffentlicht: (2024)
CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
von: Sarkar, Sayan Deb, et al.
Veröffentlicht: (2026)
von: Sarkar, Sayan Deb, et al.
Veröffentlicht: (2026)
Multi Activity Sequence Alignment via Implicit Clustering
von: Kwon, Taein, et al.
Veröffentlicht: (2025)
von: Kwon, Taein, et al.
Veröffentlicht: (2025)
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
von: Fedele, Elisabetta, et al.
Veröffentlicht: (2025)
von: Fedele, Elisabetta, et al.
Veröffentlicht: (2025)
MR.NAVI: Mixed-Reality Navigation Assistant for the Visually Impaired
von: Pfitzer, Nicolas, et al.
Veröffentlicht: (2025)
von: Pfitzer, Nicolas, et al.
Veröffentlicht: (2025)
EgoGen: An Egocentric Synthetic Data Generator
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
von: Huang, Zhao, et al.
Veröffentlicht: (2025)
von: Huang, Zhao, et al.
Veröffentlicht: (2025)
OpenDAS: Open-Vocabulary Domain Adaptation for 2D and 3D Segmentation
von: Yilmaz, Gonca, et al.
Veröffentlicht: (2024)
von: Yilmaz, Gonca, et al.
Veröffentlicht: (2024)
Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2025)
Volumetric Semantically Consistent 3D Panoptic Mapping
von: Miao, Yang, et al.
Veröffentlicht: (2023)
von: Miao, Yang, et al.
Veröffentlicht: (2023)
WildGaussians: 3D Gaussian Splatting in the Wild
von: Kulhanek, Jonas, et al.
Veröffentlicht: (2024)
von: Kulhanek, Jonas, et al.
Veröffentlicht: (2024)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
von: Wang, Pan, et al.
Veröffentlicht: (2025)
von: Wang, Pan, et al.
Veröffentlicht: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
SuperDec: 3D Scene Decomposition with Superquadric Primitives
von: Fedele, Elisabetta, et al.
Veröffentlicht: (2025)
von: Fedele, Elisabetta, et al.
Veröffentlicht: (2025)
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
von: Di Lorenzo, Gaia, et al.
Veröffentlicht: (2025)
von: Di Lorenzo, Gaia, et al.
Veröffentlicht: (2025)
3D Question Answering via only 2D Vision-Language Models
von: Wang, Fengyun, et al.
Veröffentlicht: (2025)
von: Wang, Fengyun, et al.
Veröffentlicht: (2025)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
von: Hu, Xinggang, et al.
Veröffentlicht: (2026)
von: Hu, Xinggang, et al.
Veröffentlicht: (2026)
CrossOver: 3D Scene Cross-Modal Alignment
von: Sarkar, Sayan Deb, et al.
Veröffentlicht: (2025)
von: Sarkar, Sayan Deb, et al.
Veröffentlicht: (2025)
3D Neural Edge Reconstruction
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
TORA: Topological Representation Alignment for 3D Shape Assembly
von: Lee, Nahyuk, et al.
Veröffentlicht: (2026)
von: Lee, Nahyuk, et al.
Veröffentlicht: (2026)
Dynamic 3D Gaussian Fields for Urban Areas
von: Fischer, Tobias, et al.
Veröffentlicht: (2024)
von: Fischer, Tobias, et al.
Veröffentlicht: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
von: Ye, Botao, et al.
Veröffentlicht: (2025)
von: Ye, Botao, et al.
Veröffentlicht: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
von: Yang, Yung-Hsu, et al.
Veröffentlicht: (2025)
von: Yang, Yung-Hsu, et al.
Veröffentlicht: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
Sat2Scene: 3D Urban Scene Generation from Satellite Images with Diffusion
von: Li, Zuoyue, et al.
Veröffentlicht: (2024)
von: Li, Zuoyue, et al.
Veröffentlicht: (2024)
Synthesizing Consistent Novel Views via 3D Epipolar Attention without Re-Training
von: Ye, Botao, et al.
Veröffentlicht: (2025)
von: Ye, Botao, et al.
Veröffentlicht: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
von: Li, Wenli, et al.
Veröffentlicht: (2026)
von: Li, Wenli, et al.
Veröffentlicht: (2026)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
von: Qu, Kevin, et al.
Veröffentlicht: (2026) -
MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References
von: Bösiger, Lukas, et al.
Veröffentlicht: (2024) -
CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
von: Sarkar, Sayan Deb, et al.
Veröffentlicht: (2026) -
Multi Activity Sequence Alignment via Implicit Clustering
von: Kwon, Taein, et al.
Veröffentlicht: (2025) -
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
von: Fedele, Elisabetta, et al.
Veröffentlicht: (2025)