Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yang, He, Tian, Fu, Junfeng, Wang, Ling, Guo, Jingcai, Hu, Ting, Cheng, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
by: Huang, Shunyu, et al.
Published: (2026)
by: Huang, Shunyu, et al.
Published: (2026)
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
by: Hu, Jiayun, et al.
Published: (2025)
by: Hu, Jiayun, et al.
Published: (2025)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
by: Cui, Yang, et al.
Published: (2025)
by: Cui, Yang, et al.
Published: (2025)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
by: Zhu, Xiaofei, et al.
Published: (2024)
by: Zhu, Xiaofei, et al.
Published: (2024)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
by: Fei, Ben, et al.
Published: (2024)
by: Fei, Ben, et al.
Published: (2024)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
by: An, Xiao, et al.
Published: (2026)
by: An, Xiao, et al.
Published: (2026)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
by: Lee, SangEun, et al.
Published: (2025)
by: Lee, SangEun, et al.
Published: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Rethinking Multi-view Representation Learning via Distilled Disentangling
by: Ke, Guanzhou, et al.
Published: (2024)
by: Ke, Guanzhou, et al.
Published: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
by: Sudarsanam, Parthasaarathy, et al.
Published: (2025)
by: Sudarsanam, Parthasaarathy, et al.
Published: (2025)
Contrastive Knowledge Distillation for Robust Multimodal Sentiment Analysis
by: Sang, Zhongyi, et al.
Published: (2024)
by: Sang, Zhongyi, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)
by: Gong, Ziyu, et al.
Published: (2024)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
by: Wang, Haoxu, et al.
Published: (2024)
by: Wang, Haoxu, et al.
Published: (2024)
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
by: Xiong, Zechang, et al.
Published: (2026)
by: Xiong, Zechang, et al.
Published: (2026)
Enhancing Cross-Prompt Transferability in Vision-Language Models through Contextual Injection of Target Tokens
by: Yang, Xikang, et al.
Published: (2024)
by: Yang, Xikang, et al.
Published: (2024)
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
by: Shen, Meng, et al.
Published: (2024)
by: Shen, Meng, et al.
Published: (2024)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
by: Nguyen, Manh Duong, et al.
Published: (2024)
by: Nguyen, Manh Duong, et al.
Published: (2024)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026)
by: Xu, Siyuan, et al.
Published: (2026)
Multimodal Representation Learning and Fusion
by: Jin, Qihang, et al.
Published: (2025)
by: Jin, Qihang, et al.
Published: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
by: Zhang, Bo, et al.
Published: (2024)
by: Zhang, Bo, et al.
Published: (2024)
Rethinking Fusion: Disentangled Learning of Shared and Modality-Specific Information for Stance Detection
by: Xie, Zhiyu, et al.
Published: (2026)
by: Xie, Zhiyu, et al.
Published: (2026)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
by: Wang, Yabing, et al.
Published: (2023)
by: Wang, Yabing, et al.
Published: (2023)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
by: Chang, Haochen, et al.
Published: (2024)
by: Chang, Haochen, et al.
Published: (2024)
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
by: Tang, Hao, et al.
Published: (2026)
by: Tang, Hao, et al.
Published: (2026)
Similar Items
-
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
by: Huang, Shunyu, et al.
Published: (2026) -
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
by: Hu, Jiayun, et al.
Published: (2025) -
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025) -
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
by: Zhao, Yu, et al.
Published: (2025) -
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
by: Xu, Yang, et al.
Published: (2024)