Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Ye, Luo, Weixun, Huang, Ranran, Jing, Junpeng, Mikolajczyk, Krystian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POMA-3D: The Point Map Way to 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2025)
by: Mao, Ye, et al.
Published: (2025)
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
by: Huang, Ranran, et al.
Published: (2026)
by: Huang, Ranran, et al.
Published: (2026)
Hypo3D: Exploring Hypothetical Reasoning in 3D
by: Mao, Ye, et al.
Published: (2025)
by: Mao, Ye, et al.
Published: (2025)
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
by: Mao, Ye, et al.
Published: (2024)
by: Mao, Ye, et al.
Published: (2024)
No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
by: Huang, Ranran, et al.
Published: (2025)
by: Huang, Ranran, et al.
Published: (2025)
SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
by: Huang, Ranran, et al.
Published: (2025)
by: Huang, Ranran, et al.
Published: (2025)
Match-Stereo-Videos: Bidirectional Alignment for Consistent Dynamic Stereo Matching
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Match Stereo Videos via Bidirectional Alignment
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Understanding the Role of the Projector in Knowledge Distillation
by: Miles, Roy, et al.
Published: (2023)
by: Miles, Roy, et al.
Published: (2023)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
Language-Based Depth Hints for Monocular Depth Estimation
by: Auty, Dylan, et al.
Published: (2024)
by: Auty, Dylan, et al.
Published: (2024)
UCorr: Wire Detection and Depth Estimation for Autonomous Drones
by: Kolbeinsson, Benedikt, et al.
Published: (2025)
by: Kolbeinsson, Benedikt, et al.
Published: (2025)
Multi-Class Segmentation from Aerial Views using Recursive Noise Diffusion
by: Kolbeinsson, Benedikt, et al.
Published: (2022)
by: Kolbeinsson, Benedikt, et al.
Published: (2022)
DDOS: The Drone Depth and Obstacle Segmentation Dataset
by: Kolbeinsson, Benedikt, et al.
Published: (2023)
by: Kolbeinsson, Benedikt, et al.
Published: (2023)
Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?
by: Maheshwari, Animesh, et al.
Published: (2026)
by: Maheshwari, Animesh, et al.
Published: (2026)
Contrastive Gaussian Clustering: Weakly Supervised 3D Scene Segmentation
by: Silva, Myrna C., et al.
Published: (2024)
by: Silva, Myrna C., et al.
Published: (2024)
ESGNN: Towards Equivariant Scene Graph Neural Network for 3D Scene Understanding
by: Pham, Quang P. M., et al.
Published: (2024)
by: Pham, Quang P. M., et al.
Published: (2024)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
Dynamic Scene Understanding from Vision-Language Representations
by: Pruss, Shahaf, et al.
Published: (2025)
by: Pruss, Shahaf, et al.
Published: (2025)
OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
TESGNN: Temporal Equivariant Scene Graph Neural Networks for Efficient and Robust Multi-View 3D Scene Understanding
by: Pham, Quang P. M., et al.
Published: (2024)
by: Pham, Quang P. M., et al.
Published: (2024)
Unifying Graph Contrastive Learning via Graph Message Augmentation
by: Zhang, Ziyan, et al.
Published: (2024)
by: Zhang, Ziyan, et al.
Published: (2024)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
by: Jia, Baoxiong, et al.
Published: (2024)
by: Jia, Baoxiong, et al.
Published: (2024)
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
by: Liu, I-Chun Arthur, et al.
Published: (2026)
by: Liu, I-Chun Arthur, et al.
Published: (2026)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
by: Ahmed, Fatmaelzahraa Ali, et al.
Published: (2025)
by: Ahmed, Fatmaelzahraa Ali, et al.
Published: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
by: Jiang, Minchao, et al.
Published: (2025)
by: Jiang, Minchao, et al.
Published: (2025)
Time-to-Event Pretraining for 3D Medical Imaging
by: Huo, Zepeng, et al.
Published: (2024)
by: Huo, Zepeng, et al.
Published: (2024)
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene Understanding
by: Chou, Zi-Ting, et al.
Published: (2024)
by: Chou, Zi-Ting, et al.
Published: (2024)
H2VLR: Heterogeneous Hypergraph Vision-Language Reasoning for Few-Shot Anomaly Detection
by: Huang, Jianghong, et al.
Published: (2026)
by: Huang, Jianghong, et al.
Published: (2026)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
Similar Items
-
POMA-3D: The Point Map Way to 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2025) -
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
by: Huang, Ranran, et al.
Published: (2026) -
Hypo3D: Exploring Hypothetical Reasoning in 3D
by: Mao, Ye, et al.
Published: (2025) -
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025) -
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)