RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jihan, Ding, Runyu, Deng, Weipeng, Wang, Zhe, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can 3D Vision-Language Models Truly Understand Natural Language?
by: Deng, Weipeng, et al.
Published: (2024)
by: Deng, Weipeng, et al.
Published: (2024)
V-IRL: Grounding Virtual Intelligence in Real Life
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision
by: Wang, Yuru, et al.
Published: (2024)
by: Wang, Yuru, et al.
Published: (2024)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
by: Mohiuddin, Rafay, et al.
Published: (2024)
by: Mohiuddin, Rafay, et al.
Published: (2024)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
by: Yang, Hongcheng, et al.
Published: (2024)
by: Yang, Hongcheng, et al.
Published: (2024)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
Total-Decom: Decomposed 3D Scene Reconstruction with Minimal Interaction
by: Lyu, Xiaoyang, et al.
Published: (2024)
by: Lyu, Xiaoyang, et al.
Published: (2024)
Scene Understanding Enabled Semantic Communication with Open Channel Coding
by: Xiang, Zhe, et al.
Published: (2025)
by: Xiang, Zhe, et al.
Published: (2025)
GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
by: Wang, Xihan, et al.
Published: (2025)
by: Wang, Xihan, et al.
Published: (2025)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
POMA-3D: The Point Map Way to 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2025)
by: Mao, Ye, et al.
Published: (2025)
Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding
by: Yang, Yu-Qi, et al.
Published: (2024)
by: Yang, Yu-Qi, et al.
Published: (2024)
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2026)
by: Mao, Ye, et al.
Published: (2026)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
by: Zhou, Shengchao, et al.
Published: (2025)
by: Zhou, Shengchao, et al.
Published: (2025)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
by: Zhou, Shengchao, et al.
Published: (2025)
by: Zhou, Shengchao, et al.
Published: (2025)
Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
by: Yang, Bin, et al.
Published: (2026)
by: Yang, Bin, et al.
Published: (2026)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
by: Ma, Chuofan, et al.
Published: (2025)
by: Ma, Chuofan, et al.
Published: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
by: Chandhok, Shivam
Published: (2024)
by: Chandhok, Shivam
Published: (2024)
QDM: Quadtree-Based Region-Adaptive Sparse Diffusion Models for Efficient Image Super-Resolution
by: Yang, Donglin, et al.
Published: (2025)
by: Yang, Donglin, et al.
Published: (2025)
RegionGPT: Towards Region Understanding Vision Language Model
by: Guo, Qiushan, et al.
Published: (2024)
by: Guo, Qiushan, et al.
Published: (2024)
OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding
by: Zhao, Youjun, et al.
Published: (2024)
by: Zhao, Youjun, et al.
Published: (2024)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
Unified 3D Scene Understanding Through Physical World Modeling
by: Lee, Wanhee, et al.
Published: (2026)
by: Lee, Wanhee, et al.
Published: (2026)
Parameter-efficient Prompt Learning for 3D Point Cloud Understanding
by: Sun, Hongyu, et al.
Published: (2024)
by: Sun, Hongyu, et al.
Published: (2024)
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
OpenSGA: Efficient 3D Scene Graph Alignment in the Open World
by: Chen, Gang, et al.
Published: (2026)
by: Chen, Gang, et al.
Published: (2026)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
by: Mao, Ye, et al.
Published: (2024)
by: Mao, Ye, et al.
Published: (2024)
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
by: Wu, Yanmin, et al.
Published: (2024)
by: Wu, Yanmin, et al.
Published: (2024)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes
by: Yang, Yixuan, et al.
Published: (2025)
by: Yang, Yixuan, et al.
Published: (2025)
VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions
by: Lu, Haoang, et al.
Published: (2025)
by: Lu, Haoang, et al.
Published: (2025)
Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning
by: Ding, Runyu, et al.
Published: (2024)
by: Ding, Runyu, et al.
Published: (2024)
Similar Items
-
Can 3D Vision-Language Models Truly Understand Natural Language?
by: Deng, Weipeng, et al.
Published: (2024) -
V-IRL: Grounding Virtual Intelligence in Real Life
by: Yang, Jihan, et al.
Published: (2024) -
UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision
by: Wang, Yuru, et al.
Published: (2024) -
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025) -
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
by: Mohiuddin, Rafay, et al.
Published: (2024)