Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Zuo, Zhi, Zhuang, Chenyi, Gao, Pan, Qin, Jie, Feng, Hao, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
by: Zhang, Lirui, et al.
Published: (2025)
by: Zhang, Lirui, et al.
Published: (2025)
Loomis Painter: Reconstructing the Painting Process
by: Pobitzer, Markus, et al.
Published: (2025)
by: Pobitzer, Markus, et al.
Published: (2025)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
by: Xu, Xiaoxu, et al.
Published: (2024)
by: Xu, Xiaoxu, et al.
Published: (2024)
CDFormer:When Degradation Prediction Embraces Diffusion Model for Blind Image Super-Resolution
by: Liu, Qingguo, et al.
Published: (2024)
by: Liu, Qingguo, et al.
Published: (2024)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function
by: Zhuang, Chenyi, et al.
Published: (2024)
by: Zhuang, Chenyi, et al.
Published: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
by: Xue, Feng, et al.
Published: (2025)
by: Xue, Feng, et al.
Published: (2025)
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
by: Mei, Guofeng, et al.
Published: (2025)
by: Mei, Guofeng, et al.
Published: (2025)
LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning
by: Ohamouddou, Said, et al.
Published: (2026)
by: Ohamouddou, Said, et al.
Published: (2026)
Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
by: Lu, Jingjing, et al.
Published: (2025)
by: Lu, Jingjing, et al.
Published: (2025)
Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric Augmentation
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
by: Knaebel, Karim, et al.
Published: (2023)
by: Knaebel, Karim, et al.
Published: (2023)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
by: Yu, Xinxing, et al.
Published: (2026)
by: Yu, Xinxing, et al.
Published: (2026)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
by: Han, Ruiyan, et al.
Published: (2026)
by: Han, Ruiyan, et al.
Published: (2026)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
by: Xie, Chang, et al.
Published: (2025)
by: Xie, Chang, et al.
Published: (2025)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
A Unified Framework for Human-centric Point Cloud Video Understanding
by: Xu, Yiteng, et al.
Published: (2024)
by: Xu, Yiteng, et al.
Published: (2024)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
by: Ye, Weixin, et al.
Published: (2026)
by: Ye, Weixin, et al.
Published: (2026)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
Video-Browser: Towards Agentic Open-web Video Browsing
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
by: Hu, Ying, et al.
Published: (2024)
by: Hu, Ying, et al.
Published: (2024)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024)
by: Liu, Dingning, et al.
Published: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Similar Items
-
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025) -
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
by: Ren, Bin, et al.
Published: (2024) -
Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
by: Zhang, Lirui, et al.
Published: (2025) -
Loomis Painter: Reconstructing the Painting Process
by: Pobitzer, Markus, et al.
Published: (2025) -
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)