Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yujia, Wu, Xiaoyang, Lao, Yixing, Wang, Chengyao, Tian, Zhuotao, Wang, Naiyan, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
by: Wang, Chengyao, et al.
Published: (2024)
by: Wang, Chengyao, et al.
Published: (2024)
Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
by: Wu, Xiaoyang, et al.
Published: (2023)
by: Wu, Xiaoyang, et al.
Published: (2023)
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
Pixel-GS: Density Control with Pixel-aware Gradient for 3D Gaussian Splatting
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
Utonia: Toward One Encoder for All Point Clouds
by: Zhang, Yujia, et al.
Published: (2026)
by: Zhang, Yujia, et al.
Published: (2026)
Sonata: Self-Supervised Learning of Reliable Point Representations
by: Wu, Xiaoyang, et al.
Published: (2025)
by: Wu, Xiaoyang, et al.
Published: (2025)
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
OpenIns3D: Snap and Lookup for 3D Open-vocabulary Instance Segmentation
by: Huang, Zhening, et al.
Published: (2023)
by: Huang, Zhening, et al.
Published: (2023)
3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning
by: Hu, Naiwen, et al.
Published: (2024)
by: Hu, Naiwen, et al.
Published: (2024)
Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
by: Lao, Yixing, et al.
Published: (2026)
by: Lao, Yixing, et al.
Published: (2026)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
by: Qi, Zhangyang, et al.
Published: (2024)
by: Qi, Zhangyang, et al.
Published: (2024)
Efficient 3D Perception on Multi-Sweep Point Cloud with Gumbel Spatial Pruning
by: Sun, Tianyu, et al.
Published: (2024)
by: Sun, Tianyu, et al.
Published: (2024)
LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
by: Huang, Zhening, et al.
Published: (2025)
by: Huang, Zhening, et al.
Published: (2025)
Self-Supervised Representation Learning with Spatial-Temporal Consistency for Sign Language Recognition
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Object Concepts Emerge from Motion
by: Liang, Haoqian, et al.
Published: (2025)
by: Liang, Haoqian, et al.
Published: (2025)
Enhancing 3D Lane Detection and Topology Reasoning with 2D Lane Priors
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
by: Leijenaar, Remco F., et al.
Published: (2025)
by: Leijenaar, Remco F., et al.
Published: (2025)
Cross-Dimensional Medical Self-Supervised Representation Learning Based on a Pseudo-3D Transformation
by: Gao, Fei, et al.
Published: (2024)
by: Gao, Fei, et al.
Published: (2024)
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
by: Zheng, Chaoda, et al.
Published: (2024)
by: Zheng, Chaoda, et al.
Published: (2024)
GIFS: Neural Implicit Function for General Shape Representation
by: Ye, Jianglong, et al.
Published: (2022)
by: Ye, Jianglong, et al.
Published: (2022)
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
by: Wang, Zitian, et al.
Published: (2024)
by: Wang, Zitian, et al.
Published: (2024)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
by: Gao, Yunhe, et al.
Published: (2026)
by: Gao, Yunhe, et al.
Published: (2026)
Geometry-Guided 3D Visual Token Pruning for Video-Language Models
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models
by: Ye, Jianglong, et al.
Published: (2023)
by: Ye, Jianglong, et al.
Published: (2023)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
by: Tang, Longxiang, et al.
Published: (2024)
by: Tang, Longxiang, et al.
Published: (2024)
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor Regression
by: Huang, Shaofei, et al.
Published: (2024)
by: Huang, Shaofei, et al.
Published: (2024)
SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Self-Supervised Representation Learning for Nerve Fiber Distribution Patterns in 3D-PLI
by: Oberstrass, Alexander, et al.
Published: (2024)
by: Oberstrass, Alexander, et al.
Published: (2024)
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Edit360: 2D Image Edits to 3D Assets from Any Angle
by: Huang, Junchao, et al.
Published: (2025)
by: Huang, Junchao, et al.
Published: (2025)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
Weakly Supervised Spatial Implicit Neural Representation Learning for 3D MRI-Ultrasound Deformable Image Registration in HDR Prostate Brachytherapy
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
by: Wu, Xiuzhe, et al.
Published: (2024)
by: Wu, Xiuzhe, et al.
Published: (2024)
Fully Sparse Fusion for 3D Object Detection
by: Li, Yingyan, et al.
Published: (2023)
by: Li, Yingyan, et al.
Published: (2023)
GeoDiff3D: Self-Supervised 3D Scene Generation with Geometry-Constrained 2D Diffusion Guidance
by: Zhu, Haozhi, et al.
Published: (2026)
by: Zhu, Haozhi, et al.
Published: (2026)
Similar Items
-
GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
by: Wang, Chengyao, et al.
Published: (2024) -
Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
by: Wu, Xiaoyang, et al.
Published: (2023) -
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
by: Peng, Bohao, et al.
Published: (2024) -
Pixel-GS: Density Control with Pixel-aware Gradient for 3D Gaussian Splatting
by: Zhang, Zheng, et al.
Published: (2024) -
Utonia: Toward One Encoder for All Point Clouds
by: Zhang, Yujia, et al.
Published: (2026)