Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bing, Chen, Jiaxin, Zhang, Dongming, Bao, Xiuguo, Huang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
by: Xu, Xuancheng, et al.
Published: (2025)
by: Xu, Xuancheng, et al.
Published: (2025)
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
by: Wan, Xinhang, et al.
Published: (2025)
by: Wan, Xinhang, et al.
Published: (2025)
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
by: Guo, Hanyu, et al.
Published: (2024)
by: Guo, Hanyu, et al.
Published: (2024)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
Learning Representations for Clustering via Partial Information Discrimination and Cross-Level Interaction
by: Zhang, Hai-Xin, et al.
Published: (2024)
by: Zhang, Hai-Xin, et al.
Published: (2024)
Towards Scalable Modeling of Compressed Videos for Efficient Action Recognition
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
by: Qu, Mingcheng, et al.
Published: (2025)
by: Qu, Mingcheng, et al.
Published: (2025)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Learning Human Motion from Monocular Videos via Cross-Modal Manifold Alignment
by: Hou, Shuaiying, et al.
Published: (2024)
by: Hou, Shuaiying, et al.
Published: (2024)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
by: Huang, Zhanbo, et al.
Published: (2025)
by: Huang, Zhanbo, et al.
Published: (2025)
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
by: Chen, Jiajun, et al.
Published: (2025)
by: Chen, Jiajun, et al.
Published: (2025)
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
by: Zheng, Meng, et al.
Published: (2024)
by: Zheng, Meng, et al.
Published: (2024)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
An Interpretable Cross-Attentive Multi-modal MRI Fusion Framework for Schizophrenia Diagnosis
by: Zhou, Ziyu, et al.
Published: (2024)
by: Zhou, Ziyu, et al.
Published: (2024)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
iVPT: Improving Task-relevant Information Sharing in Visual Prompt Tuning by Cross-layer Dynamic Connection
by: Zhou, Nan, et al.
Published: (2024)
by: Zhou, Nan, et al.
Published: (2024)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
A Self-supervised Motion Representation for Portrait Video Generation
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
Trusted Video Inpainting Localization via Deep Attentive Noise Learning
by: Lou, Zijie, et al.
Published: (2024)
by: Lou, Zijie, et al.
Published: (2024)
SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition
by: Peng, Zhuoxuan, et al.
Published: (2026)
by: Peng, Zhuoxuan, et al.
Published: (2026)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025)
by: He, Wen-Jue, et al.
Published: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
by: Hu, Yunzuo, et al.
Published: (2026)
by: Hu, Yunzuo, et al.
Published: (2026)
Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
by: Yuan, Junyi, et al.
Published: (2025)
by: Yuan, Junyi, et al.
Published: (2025)
Advancing Compressed Video Action Recognition through Progressive Knowledge Distillation
by: Soufleri, Efstathia, et al.
Published: (2024)
by: Soufleri, Efstathia, et al.
Published: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
Plug-and-Play Versatile Compressed Video Enhancement
by: Zeng, Huimin, et al.
Published: (2025)
by: Zeng, Huimin, et al.
Published: (2025)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
Compressed Video Quality Enhancement with Temporal Group Alignment and Fusion
by: Zhu, Qiang, et al.
Published: (2024)
by: Zhu, Qiang, et al.
Published: (2024)
StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain Analysis
by: Chen, Lixin, et al.
Published: (2025)
by: Chen, Lixin, et al.
Published: (2025)
DrFER: Learning Disentangled Representations for 3D Facial Expression Recognition
by: Li, Hebeizi, et al.
Published: (2024)
by: Li, Hebeizi, et al.
Published: (2024)
Similar Items
-
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024) -
SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
by: Xu, Xuancheng, et al.
Published: (2025) -
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025) -
Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
by: Wan, Xinhang, et al.
Published: (2025) -
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
by: Guo, Hanyu, et al.
Published: (2024)