Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Zuo, Zhi, Zhuang, Chenyi, Gao, Pan, Qin, Jie, Feng, Hao, Sebe, Nicu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
por: Wang, Yidan, et al.
Publicado: (2025)
por: Wang, Yidan, et al.
Publicado: (2025)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
por: Ren, Bin, et al.
Publicado: (2024)
por: Ren, Bin, et al.
Publicado: (2024)
Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
por: Zhang, Lirui, et al.
Publicado: (2025)
por: Zhang, Lirui, et al.
Publicado: (2025)
Loomis Painter: Reconstructing the Painting Process
por: Pobitzer, Markus, et al.
Publicado: (2025)
por: Pobitzer, Markus, et al.
Publicado: (2025)
Fully-Geometric Cross-Attention for Point Cloud Registration
por: Wang, Weijie, et al.
Publicado: (2025)
por: Wang, Weijie, et al.
Publicado: (2025)
Asymmetric GANs for Image-to-Image Translation
por: Tang, Hao, et al.
Publicado: (2019)
por: Tang, Hao, et al.
Publicado: (2019)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
por: Ren, Bin, et al.
Publicado: (2025)
por: Ren, Bin, et al.
Publicado: (2025)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
por: Pang, Youxin, et al.
Publicado: (2025)
por: Pang, Youxin, et al.
Publicado: (2025)
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
por: Xu, Xiaoxu, et al.
Publicado: (2024)
por: Xu, Xiaoxu, et al.
Publicado: (2024)
CDFormer:When Degradation Prediction Embraces Diffusion Model for Blind Image Super-Resolution
por: Liu, Qingguo, et al.
Publicado: (2024)
por: Liu, Qingguo, et al.
Publicado: (2024)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
por: Li, Jinlong, et al.
Publicado: (2024)
por: Li, Jinlong, et al.
Publicado: (2024)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
por: Wang, Weijie, et al.
Publicado: (2023)
por: Wang, Weijie, et al.
Publicado: (2023)
Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function
por: Zhuang, Chenyi, et al.
Publicado: (2024)
por: Zhuang, Chenyi, et al.
Publicado: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
por: Li, Jinlong, et al.
Publicado: (2026)
por: Li, Jinlong, et al.
Publicado: (2026)
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
por: Xue, Feng, et al.
Publicado: (2025)
por: Xue, Feng, et al.
Publicado: (2025)
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
por: Mei, Guofeng, et al.
Publicado: (2025)
por: Mei, Guofeng, et al.
Publicado: (2025)
LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning
por: Ohamouddou, Said, et al.
Publicado: (2026)
por: Ohamouddou, Said, et al.
Publicado: (2026)
Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
por: Lu, Jingjing, et al.
Publicado: (2025)
por: Lu, Jingjing, et al.
Publicado: (2025)
Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric Augmentation
por: Yu, Li, et al.
Publicado: (2024)
por: Yu, Li, et al.
Publicado: (2024)
UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
por: Knaebel, Karim, et al.
Publicado: (2023)
por: Knaebel, Karim, et al.
Publicado: (2023)
Hierarchical Cross-Attention Network for Virtual Try-On
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
por: Yu, Xinxing, et al.
Publicado: (2026)
por: Yu, Xinxing, et al.
Publicado: (2026)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
por: Han, Ruiyan, et al.
Publicado: (2026)
por: Han, Ruiyan, et al.
Publicado: (2026)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
por: Xie, Chang, et al.
Publicado: (2025)
por: Xie, Chang, et al.
Publicado: (2025)
Rethinking the Learning Paradigm for Facial Expression Recognition
por: Wang, Weijie, et al.
Publicado: (2022)
por: Wang, Weijie, et al.
Publicado: (2022)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
por: Li, Jinlong, et al.
Publicado: (2025)
por: Li, Jinlong, et al.
Publicado: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
por: Xing, Songlong, et al.
Publicado: (2025)
por: Xing, Songlong, et al.
Publicado: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
por: Mao, Weijia, et al.
Publicado: (2025)
por: Mao, Weijia, et al.
Publicado: (2025)
A Unified Framework for Human-centric Point Cloud Video Understanding
por: Xu, Yiteng, et al.
Publicado: (2024)
por: Xu, Yiteng, et al.
Publicado: (2024)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
por: Li, Yidi, et al.
Publicado: (2024)
por: Li, Yidi, et al.
Publicado: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
por: Peruzzo, Elia, et al.
Publicado: (2025)
por: Peruzzo, Elia, et al.
Publicado: (2025)
Vision+X: A Survey on Multimodal Learning in the Light of Data
por: Zhu, Ye, et al.
Publicado: (2022)
por: Zhu, Ye, et al.
Publicado: (2022)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
por: Ye, Weixin, et al.
Publicado: (2026)
por: Ye, Weixin, et al.
Publicado: (2026)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
por: Sacilotti, André, et al.
Publicado: (2024)
por: Sacilotti, André, et al.
Publicado: (2024)
Video-Browser: Towards Agentic Open-web Video Browsing
por: Liang, Zhengyang, et al.
Publicado: (2025)
por: Liang, Zhengyang, et al.
Publicado: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
por: Hu, Ying, et al.
Publicado: (2024)
por: Hu, Ying, et al.
Publicado: (2024)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
por: Liu, Dingning, et al.
Publicado: (2024)
por: Liu, Dingning, et al.
Publicado: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
Ejemplares similares
-
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
por: Wang, Yidan, et al.
Publicado: (2025) -
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
por: Ren, Bin, et al.
Publicado: (2024) -
Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
por: Zhang, Lirui, et al.
Publicado: (2025) -
Loomis Painter: Reconstructing the Painting Process
por: Pobitzer, Markus, et al.
Publicado: (2025) -
Fully-Geometric Cross-Attention for Point Cloud Registration
por: Wang, Weijie, et al.
Publicado: (2025)