From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Tianxu, Zhang, Zhuofan, Zhu, Ziyu, Fan, Yue, Xiong, Jing, Li, Pengxiang, Ma, Xiaojian, Li, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Task-oriented Sequential Grounding and Navigation in 3D Scenes
von: Zhang, Zhuofan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhuofan, et al.
Veröffentlicht: (2024)
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
von: Zhu, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyu, et al.
Veröffentlicht: (2025)
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
von: Guo, Jun, et al.
Veröffentlicht: (2024)
von: Guo, Jun, et al.
Veröffentlicht: (2024)
Unifying 3D Vision-Language Understanding via Promptable Queries
von: Zhu, Ziyu, et al.
Veröffentlicht: (2024)
von: Zhu, Ziyu, et al.
Veröffentlicht: (2024)
DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
von: Dominici, Edoardo Alberto, et al.
Veröffentlicht: (2025)
von: Dominici, Edoardo Alberto, et al.
Veröffentlicht: (2025)
Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2025)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2025)
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
von: Gao, Zhi, et al.
Veröffentlicht: (2024)
von: Gao, Zhi, et al.
Veröffentlicht: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond
von: Zhang, Haiming, et al.
Veröffentlicht: (2026)
von: Zhang, Haiming, et al.
Veröffentlicht: (2026)
Multi-modal Situated Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
von: Jin, Hoiyeong, et al.
Veröffentlicht: (2025)
von: Jin, Hoiyeong, et al.
Veröffentlicht: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
von: Hu, Miao, et al.
Veröffentlicht: (2025)
von: Hu, Miao, et al.
Veröffentlicht: (2025)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
von: Wu, Zehuan, et al.
Veröffentlicht: (2024)
von: Wu, Zehuan, et al.
Veröffentlicht: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
von: Lu, Jialin, et al.
Veröffentlicht: (2024)
von: Lu, Jialin, et al.
Veröffentlicht: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
von: Lu, Jialin, et al.
Veröffentlicht: (2025)
von: Lu, Jialin, et al.
Veröffentlicht: (2025)
GauU-Scene: A Scene Reconstruction Benchmark on Large Scale 3D Reconstruction Dataset Using Gaussian Splatting
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence
von: Shi, Zhuofan, et al.
Veröffentlicht: (2026)
von: Shi, Zhuofan, et al.
Veröffentlicht: (2026)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
Recent Advances in 3D Object and Scene Generation: A Survey
von: Tang, Xiang, et al.
Veröffentlicht: (2025)
von: Tang, Xiang, et al.
Veröffentlicht: (2025)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
VisAnywhere: Developing Multi-platform Scientific Visualization Applications
von: Marrinan, Thomas, et al.
Veröffentlicht: (2024)
von: Marrinan, Thomas, et al.
Veröffentlicht: (2024)
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
von: Zhou, Hongyu, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyu, et al.
Veröffentlicht: (2024)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
von: Zuo, Xingxing, et al.
Veröffentlicht: (2024)
von: Zuo, Xingxing, et al.
Veröffentlicht: (2024)
RoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware Roaming
von: Chu, Jisheng, et al.
Veröffentlicht: (2026)
von: Chu, Jisheng, et al.
Veröffentlicht: (2026)
WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language
von: Lin, Zhenxiang, et al.
Veröffentlicht: (2023)
von: Lin, Zhenxiang, et al.
Veröffentlicht: (2023)
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2024)
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2024)
Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
von: Zhang, Jingbo, et al.
Veröffentlicht: (2023)
von: Zhang, Jingbo, et al.
Veröffentlicht: (2023)
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
von: Wu, Guile, et al.
Veröffentlicht: (2026)
von: Wu, Guile, et al.
Veröffentlicht: (2026)
Multi-Task Domain Adaptation for Language Grounding with 3D Objects
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
Benchmarking Single-Step Inpainting Methods for Multi-Object 3D Gaussian Splatting Scenes
von: Dröge, Finn, et al.
Veröffentlicht: (2026)
von: Dröge, Finn, et al.
Veröffentlicht: (2026)
Think Anywhere in Code Generation
von: Jiang, Xue, et al.
Veröffentlicht: (2026)
von: Jiang, Xue, et al.
Veröffentlicht: (2026)
Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Task-oriented Sequential Grounding and Navigation in 3D Scenes
von: Zhang, Zhuofan, et al.
Veröffentlicht: (2024) -
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
von: Zhu, Ziyu, et al.
Veröffentlicht: (2025) -
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
von: Guo, Jun, et al.
Veröffentlicht: (2024) -
Unifying 3D Vision-Language Understanding via Promptable Queries
von: Zhu, Ziyu, et al.
Veröffentlicht: (2024) -
DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
von: Dominici, Edoardo Alberto, et al.
Veröffentlicht: (2025)