CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yunze, Chen, Changxi, Wang, Zifan, Yi, Li |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Interactive Humanoid: Online Full-Body Motion Reaction Synthesis with Social Affordance Canonicalization and Forecasting
by: Liu, Yunze, et al.
Published: (2023)
by: Liu, Yunze, et al.
Published: (2023)
A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
by: Wu, Chengzhi, et al.
Published: (2025)
by: Wu, Chengzhi, et al.
Published: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
by: Zhang, Runmin, et al.
Published: (2024)
by: Zhang, Runmin, et al.
Published: (2024)
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
PhysReaction: Physically Plausible Real-Time Humanoid Reaction Synthesis via Forward Dynamics Guided 4D Imitation
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
by: Zha, Yixin, et al.
Published: (2025)
by: Zha, Yixin, et al.
Published: (2025)
Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion
by: Xu, Hang, et al.
Published: (2024)
by: Xu, Hang, et al.
Published: (2024)
Self-supervised Event-based Monocular Depth Estimation using Cross-modal Consistency
by: Zhu, Junyu, et al.
Published: (2024)
by: Zhu, Junyu, et al.
Published: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
GS-PT: Exploiting 3D Gaussian Splatting for Comprehensive Point Cloud Understanding via Self-supervised Learning
by: Liu, Keyi, et al.
Published: (2024)
by: Liu, Keyi, et al.
Published: (2024)
EPContrast: Effective Point-level Contrastive Learning for Large-scale Point Cloud Understanding
by: Pan, Zhiyi, et al.
Published: (2024)
by: Pan, Zhiyi, et al.
Published: (2024)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Slimmable Networks for Contrastive Self-supervised Learning
by: Zhao, Shuai, et al.
Published: (2022)
by: Zhao, Shuai, et al.
Published: (2022)
PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
by: Yu, Xinxing, et al.
Published: (2026)
by: Yu, Xinxing, et al.
Published: (2026)
Learning to Adapt SAM for Segmenting Cross-domain Point Clouds
by: Peng, Xidong, et al.
Published: (2023)
by: Peng, Xidong, et al.
Published: (2023)
PointCG: Self-supervised Point Cloud Learning via Joint Completion and Generation
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2026)
by: Wu, Peiran, et al.
Published: (2026)
Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised Learning
by: Zha, Yaohua, et al.
Published: (2024)
by: Zha, Yaohua, et al.
Published: (2024)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning
by: You, Weihang, et al.
Published: (2026)
by: You, Weihang, et al.
Published: (2026)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
by: Hou, Hongye, et al.
Published: (2025)
by: Hou, Hongye, et al.
Published: (2025)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning
by: Zhang, Yabin, et al.
Published: (2022)
by: Zhang, Yabin, et al.
Published: (2022)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
MutualNeRF: Improve the Performance of NeRF under Limited Samples with Mutual Information Theory
by: Wang, Zifan, et al.
Published: (2025)
by: Wang, Zifan, et al.
Published: (2025)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
Attention-Enhanced Cross-modal Localization Between 360 Images and Point Clouds
by: Zhao, Zhipeng, et al.
Published: (2022)
by: Zhao, Zhipeng, et al.
Published: (2022)
GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
by: Wang, Chengyao, et al.
Published: (2024)
by: Wang, Chengyao, et al.
Published: (2024)
PointDC:Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering
by: Chen, Zisheng, et al.
Published: (2023)
by: Chen, Zisheng, et al.
Published: (2023)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
by: Tao, Zhuo, et al.
Published: (2025)
by: Tao, Zhuo, et al.
Published: (2025)
Similar Items
-
PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
by: Liu, Yunze, et al.
Published: (2025) -
Interactive Humanoid: Online Full-Body Motion Reaction Synthesis with Social Affordance Canonicalization and Forecasting
by: Liu, Yunze, et al.
Published: (2023) -
A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
by: Wu, Chengzhi, et al.
Published: (2025) -
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025) -
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
by: Zhang, Runmin, et al.
Published: (2024)