Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Haifeng, Chen, Yilun, Wang, Zehan, Pang, Jiangmiao, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
Grounded 3D-LLM with Referent Tokens
von: Chen, Yilun, et al.
Veröffentlicht: (2024)
von: Chen, Yilun, et al.
Veröffentlicht: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
von: Lyu, Ruiyuan, et al.
Veröffentlicht: (2024)
von: Lyu, Ruiyuan, et al.
Veröffentlicht: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
von: Hu, Miao, et al.
Veröffentlicht: (2025)
von: Hu, Miao, et al.
Veröffentlicht: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
von: Wen, Xin, et al.
Veröffentlicht: (2025)
von: Wen, Xin, et al.
Veröffentlicht: (2025)
Multi-Object Tracking by Hierarchical Visual Representations
von: Cao, Jinkun, et al.
Veröffentlicht: (2024)
von: Cao, Jinkun, et al.
Veröffentlicht: (2024)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
von: Chen, Xiao, et al.
Veröffentlicht: (2025)
von: Chen, Xiao, et al.
Veröffentlicht: (2025)
PointLLM: Empowering Large Language Models to Understand Point Clouds
von: Xu, Runsen, et al.
Veröffentlicht: (2023)
von: Xu, Runsen, et al.
Veröffentlicht: (2023)
HaloGS: Loose Coupling of Compact Geometry and Gaussian Splats for 3D Scenes
von: Jiang, Changjian, et al.
Veröffentlicht: (2025)
von: Jiang, Changjian, et al.
Veröffentlicht: (2025)
3DGSR: Implicit Surface Reconstruction with 3D Gaussian Splatting
von: Lyu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Lyu, Xiaoyang, et al.
Veröffentlicht: (2024)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
von: Wen, Xin, et al.
Veröffentlicht: (2024)
von: Wen, Xin, et al.
Veröffentlicht: (2024)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
von: Peng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Peng, Jiaqi, et al.
Veröffentlicht: (2025)
FurnSet: Exploiting Repeats for 3D Scene Reconstruction
von: Dobre, Paul, et al.
Veröffentlicht: (2026)
von: Dobre, Paul, et al.
Veröffentlicht: (2026)
Potential Field as Scene Affordance for Behavior Change-Based Visual Risk Object Identification
von: Pao, Pang-Yuan, et al.
Veröffentlicht: (2024)
von: Pao, Pang-Yuan, et al.
Veröffentlicht: (2024)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
von: Zhu, Ruijie, et al.
Veröffentlicht: (2025)
von: Zhu, Ruijie, et al.
Veröffentlicht: (2025)
Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
von: Fang, Shuangkang, et al.
Veröffentlicht: (2024)
von: Fang, Shuangkang, et al.
Veröffentlicht: (2024)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
RoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware Roaming
von: Chu, Jisheng, et al.
Veröffentlicht: (2026)
von: Chu, Jisheng, et al.
Veröffentlicht: (2026)
Scene-Conditional 3D Object Stylization and Composition
von: Zhou, Jinghao, et al.
Veröffentlicht: (2023)
von: Zhou, Jinghao, et al.
Veröffentlicht: (2023)
Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
von: Liu, Enyu, et al.
Veröffentlicht: (2025)
von: Liu, Enyu, et al.
Veröffentlicht: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
Unified Human-Scene Interaction via Prompted Chain-of-Contacts
von: Xiao, Zeqi, et al.
Veröffentlicht: (2023)
von: Xiao, Zeqi, et al.
Veröffentlicht: (2023)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
Style-Consistent 3D Indoor Scene Synthesis with Decoupled Objects
von: Zhang, Yunfan, et al.
Veröffentlicht: (2024)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2024)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
von: Miao, Xingyu, et al.
Veröffentlicht: (2026)
von: Miao, Xingyu, et al.
Veröffentlicht: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
von: Hao, Jinkun, et al.
Veröffentlicht: (2025)
von: Hao, Jinkun, et al.
Veröffentlicht: (2025)
GenSpace: Benchmarking Spatially-Aware Image Generation
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
Improving Classification of Occluded Objects through Scene Context
von: King, Courtney M., et al.
Veröffentlicht: (2025)
von: King, Courtney M., et al.
Veröffentlicht: (2025)
IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
von: Yin, Junbo, et al.
Veröffentlicht: (2024)
von: Yin, Junbo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023) -
Grounded 3D-LLM with Referent Tokens
von: Chen, Yilun, et al.
Veröffentlicht: (2024) -
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025) -
Language-to-Space Programming for Training-Free 3D Visual Grounding
von: Mi, Boyu, et al.
Veröffentlicht: (2025) -
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)