SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Baoxiong, Chen, Yixin, Yu, Huangyue, Wang, Yan, Niu, Xuesong, Liu, Tengyu, Li, Qing, Huang, Siyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-modal Situated Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024)
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
von: Yu, Huangyue, et al.
Veröffentlicht: (2025)
von: Yu, Huangyue, et al.
Veröffentlicht: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
von: Yang, Yandan, et al.
Veröffentlicht: (2025)
von: Yang, Yandan, et al.
Veröffentlicht: (2025)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
von: Ge, Luzhou, et al.
Veröffentlicht: (2026)
Task-oriented Sequential Grounding and Navigation in 3D Scenes
von: Zhang, Zhuofan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhuofan, et al.
Veröffentlicht: (2024)
PRISM: Preference Refinement via Implicit Scene Modeling for 3D Vision-Language Preference-Based Reinforcement Learning
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
von: Sun, Yirong, et al.
Veröffentlicht: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2025)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2025)
Unifying 3D Vision-Language Understanding via Promptable Queries
von: Zhu, Ziyu, et al.
Veröffentlicht: (2024)
von: Zhu, Ziyu, et al.
Veröffentlicht: (2024)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
von: Wang, Zan, et al.
Veröffentlicht: (2024)
von: Wang, Zan, et al.
Veröffentlicht: (2024)
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
von: Zhu, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyu, et al.
Veröffentlicht: (2025)
Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
Language-Assisted 3D Scene Understanding
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
von: Cui, Jieming, et al.
Veröffentlicht: (2024)
von: Cui, Jieming, et al.
Veröffentlicht: (2024)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
von: Ma, Xianzheng, et al.
Veröffentlicht: (2026)
von: Ma, Xianzheng, et al.
Veröffentlicht: (2026)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
Active Vision for Scene Understanding
von: Grotz, Markus
Veröffentlicht: (2022)
von: Grotz, Markus
Veröffentlicht: (2022)
GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill
von: Cui, Jieming, et al.
Veröffentlicht: (2025)
von: Cui, Jieming, et al.
Veröffentlicht: (2025)
HERO: Hierarchical Traversable 3D Scene Graphs for Embodied Navigation Among Movable Obstacles
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
Stable Language Guidance for Vision-Language-Action Models
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
Embodied Agents for Efficient Exploration and Smart Scene Description
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2023)
von: Bigazzi, Roberto, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Multi-modal Situated Reasoning in 3D Scenes
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2024) -
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
von: Yu, Huangyue, et al.
Veröffentlicht: (2025) -
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
von: Wang, Yan, et al.
Veröffentlicht: (2025) -
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
von: Yang, Yandan, et al.
Veröffentlicht: (2024) -
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
von: Yang, Yandan, et al.
Veröffentlicht: (2025)