Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yifei, Chen, Dapeng, Liu, Ruijin, Zhou, Sai, Xue, Wenyuan, Peng, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Storyboard guided Alignment for Fine-grained Video Action Recognition
by: Liu, Enqi, et al.
Published: (2024)
by: Liu, Enqi, et al.
Published: (2024)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
by: Wang, Mengmeng, et al.
Published: (2024)
by: Wang, Mengmeng, et al.
Published: (2024)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
by: Yuan, Zhenlong, et al.
Published: (2025)
by: Yuan, Zhenlong, et al.
Published: (2025)
A Survey on Hallucination in Large Vision-Language Models
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts
by: Yang, Yanting, et al.
Published: (2024)
by: Yang, Yanting, et al.
Published: (2024)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
by: Nie, Dujun, et al.
Published: (2026)
by: Nie, Dujun, et al.
Published: (2026)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain Analysis
by: Chen, Lixin, et al.
Published: (2025)
by: Chen, Lixin, et al.
Published: (2025)
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps
by: Xia, Xue, et al.
Published: (2024)
by: Xia, Xue, et al.
Published: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
Denoising-Contrastive Alignment for Continuous Sign Language Recognition
by: Guo, Leming, et al.
Published: (2023)
by: Guo, Leming, et al.
Published: (2023)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
Leveraging Synthetic Data for Generalizable and Fair Facial Action Unit Detection
by: Lu, Liupei, et al.
Published: (2024)
by: Lu, Liupei, et al.
Published: (2024)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025)
by: Peng, Liyang, et al.
Published: (2025)
SkateboardAI: The Coolest Video Action Recognition for Skateboarding
by: Chen, Hanxiao
Published: (2023)
by: Chen, Hanxiao
Published: (2023)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment
by: Zhou, Kai, et al.
Published: (2025)
by: Zhou, Kai, et al.
Published: (2025)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026)
by: Zhu, Shangwen, et al.
Published: (2026)
AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model Alignment
by: Chen, Ruoxin, et al.
Published: (2025)
by: Chen, Ruoxin, et al.
Published: (2025)
Alignment-free Raw Video Demoireing
by: Xu, Shuning, et al.
Published: (2024)
by: Xu, Shuning, et al.
Published: (2024)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
by: Zhang, Yanghai, et al.
Published: (2024)
by: Zhang, Yanghai, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
by: Chen, Yin, et al.
Published: (2023)
by: Chen, Yin, et al.
Published: (2023)
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Adapt before Continual Learning
by: Lu, Aojun, et al.
Published: (2025)
by: Lu, Aojun, et al.
Published: (2025)
Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
by: Tang, Canhui, et al.
Published: (2025)
by: Tang, Canhui, et al.
Published: (2025)
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
by: Liu, Chao, et al.
Published: (2024)
by: Liu, Chao, et al.
Published: (2024)
SV3.3B: A Sports Video Understanding Model for Action Recognition
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Heatmap Pooling Network for Action Recognition from RGB Videos
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
An Animation-based Augmentation Approach for Action Recognition from Discontinuous Video
by: Song, Xingyu, et al.
Published: (2024)
by: Song, Xingyu, et al.
Published: (2024)
Similar Items
-
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024) -
Storyboard guided Alignment for Fine-grained Video Action Recognition
by: Liu, Enqi, et al.
Published: (2024) -
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
by: Wang, Mengmeng, et al.
Published: (2024) -
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024) -
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)