Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Man, Yunze, Zheng, Shuhong, Bao, Zhipeng, Hebert, Martial, Gui, Liang-Yan, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
von: Zheng, Shuhong, et al.
Veröffentlicht: (2024)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2024)
Situational Awareness Matters in 3D Vision Language Reasoning
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023)
von: Man, Yunze, et al.
Veröffentlicht: (2023)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
von: Ma, Xianzheng, et al.
Veröffentlicht: (2026)
von: Ma, Xianzheng, et al.
Veröffentlicht: (2026)
PRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
MoD-SLAM: Monocular Dense Mapping for Unbounded 3D Scene Reconstruction
von: Zhou, Heng, et al.
Veröffentlicht: (2024)
von: Zhou, Heng, et al.
Veröffentlicht: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
SceneCraft: Layout-Guided 3D Scene Generation
von: Yang, Xiuyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiuyu, et al.
Veröffentlicht: (2024)
HERO: Hierarchical Traversable 3D Scene Graphs for Embodied Navigation Among Movable Obstacles
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
Language-Assisted 3D Scene Understanding
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
3D-VLA: A 3D Vision-Language-Action Generative World Model
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
von: Saxena, Saumya, et al.
Veröffentlicht: (2024)
von: Saxena, Saumya, et al.
Veröffentlicht: (2024)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
von: Gupta, Vinayak, et al.
Veröffentlicht: (2024)
von: Gupta, Vinayak, et al.
Veröffentlicht: (2024)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
Compact 3D Gaussian Splatting For Dense Visual SLAM
von: Deng, Tianchen, et al.
Veröffentlicht: (2024)
von: Deng, Tianchen, et al.
Veröffentlicht: (2024)
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
von: Li, Kailing, et al.
Veröffentlicht: (2026)
von: Li, Kailing, et al.
Veröffentlicht: (2026)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
von: Zheng, Shuhong, et al.
Veröffentlicht: (2025)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2025)
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
von: Li, Ye, et al.
Veröffentlicht: (2024)
von: Li, Ye, et al.
Veröffentlicht: (2024)
Unifying Scene Representation and Hand-Eye Calibration with 3D Foundation Models
von: Zhi, Weiming, et al.
Veröffentlicht: (2024)
von: Zhi, Weiming, et al.
Veröffentlicht: (2024)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
Foundation Models for Autonomous Robots in Unstructured Environments
von: Naderi, Hossein, et al.
Veröffentlicht: (2024)
von: Naderi, Hossein, et al.
Veröffentlicht: (2024)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
von: Guo, Xiyue, et al.
Veröffentlicht: (2025)
von: Guo, Xiyue, et al.
Veröffentlicht: (2025)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
von: Mehan, Yash, et al.
Veröffentlicht: (2024)
von: Mehan, Yash, et al.
Veröffentlicht: (2024)
What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
von: Zheng, Shuhong, et al.
Veröffentlicht: (2024) -
Situational Awareness Matters in 3D Vision Language Reasoning
von: Man, Yunze, et al.
Veröffentlicht: (2024) -
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023) -
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024) -
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
von: Ma, Xianzheng, et al.
Veröffentlicht: (2026)