Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Haifeng, Chen, Yilun, Wang, Zehan, Huang, Rongjie, Xu, Runsen, Wang, Tai, Liu, Luping, Cheng, Xize, Zhao, Yang, Pang, Jiangmiao, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026)
by: Huang, Haifeng, et al.
Published: (2026)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
Grounded 3D-LLM with Referent Tokens
by: Chen, Yilun, et al.
Published: (2024)
by: Chen, Yilun, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025)
by: Mi, Boyu, et al.
Published: (2025)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
by: Zhu, Ruijie, et al.
Published: (2025)
by: Zhu, Ruijie, et al.
Published: (2025)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
Unified Human-Scene Interaction via Prompted Chain-of-Contacts
by: Xiao, Zeqi, et al.
Published: (2023)
by: Xiao, Zeqi, et al.
Published: (2023)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
by: Gao, Ning, et al.
Published: (2025)
by: Gao, Ning, et al.
Published: (2025)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
by: Xu, Wenting, et al.
Published: (2024)
by: Xu, Wenting, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
by: Hong, Zhiqing, et al.
Published: (2024)
by: Hong, Zhiqing, et al.
Published: (2024)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
Lightweight Visual Measurement of Tunnel Scenes Based on SAE ‐ DeepLabV3 +
by: Xuechun Shi, et al.
Published: (2025)
by: Xuechun Shi, et al.
Published: (2025)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
Toward Scene Graph and Layout Guided Complex 3D Scene Generation
by: Huang, Yu-Hsiang, et al.
Published: (2024)
by: Huang, Yu-Hsiang, et al.
Published: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
HaloGS: Loose Coupling of Compact Geometry and Gaussian Splats for 3D Scenes
by: Jiang, Changjian, et al.
Published: (2025)
by: Jiang, Changjian, et al.
Published: (2025)
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
by: Wang, Puyi, et al.
Published: (2026)
by: Wang, Puyi, et al.
Published: (2026)
S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural Field
by: Liang, Zixi, et al.
Published: (2024)
by: Liang, Zixi, et al.
Published: (2024)
Similar Items
-
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026) -
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024) -
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025) -
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024) -
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)