MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Binhua, Wang, Ni, Pakrashi, Arjun, Dev, Soumyabrata |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
di: Huang, Binhua, et al.
Pubblicazione: (2025)
di: Huang, Binhua, et al.
Pubblicazione: (2025)
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
di: Maldonado, Gabriel, et al.
Pubblicazione: (2025)
di: Maldonado, Gabriel, et al.
Pubblicazione: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
di: Huang, Binhua, et al.
Pubblicazione: (2025)
di: Huang, Binhua, et al.
Pubblicazione: (2025)
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
di: Yao, Wendong, et al.
Pubblicazione: (2025)
di: Yao, Wendong, et al.
Pubblicazione: (2025)
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
di: Yao, Wendong, et al.
Pubblicazione: (2025)
di: Yao, Wendong, et al.
Pubblicazione: (2025)
LiteEmbed: Adapting CLIP to Rare Classes
di: Agarwal, Aishwarya, et al.
Pubblicazione: (2026)
di: Agarwal, Aishwarya, et al.
Pubblicazione: (2026)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
di: Alyami, Sarah, et al.
Pubblicazione: (2025)
di: Alyami, Sarah, et al.
Pubblicazione: (2025)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
di: Ahmad, Shahzad, et al.
Pubblicazione: (2023)
di: Ahmad, Shahzad, et al.
Pubblicazione: (2023)
A Deep Learning Approach for Spatio-Temporal Forecasting of InSAR Ground Deformation in Eastern Ireland
di: Yao, Wendong, et al.
Pubblicazione: (2025)
di: Yao, Wendong, et al.
Pubblicazione: (2025)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
di: Zhu, Wencheng, et al.
Pubblicazione: (2025)
di: Zhu, Wencheng, et al.
Pubblicazione: (2025)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
di: Singh, Darshan, et al.
Pubblicazione: (2024)
di: Singh, Darshan, et al.
Pubblicazione: (2024)
AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
di: Zhong, Enmin, et al.
Pubblicazione: (2025)
di: Zhong, Enmin, et al.
Pubblicazione: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
di: Zhao, Weiheng, et al.
Pubblicazione: (2025)
di: Zhao, Weiheng, et al.
Pubblicazione: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
di: Zhang, Kangjie, et al.
Pubblicazione: (2026)
di: Zhang, Kangjie, et al.
Pubblicazione: (2026)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
di: Song, Dan, et al.
Pubblicazione: (2023)
di: Song, Dan, et al.
Pubblicazione: (2023)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
di: Zhu, Wenqi, et al.
Pubblicazione: (2024)
di: Zhu, Wenqi, et al.
Pubblicazione: (2024)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
di: Chung, Jeannie, et al.
Pubblicazione: (2026)
di: Chung, Jeannie, et al.
Pubblicazione: (2026)
Emotion Recognition with CLIP and Sequential Learning
di: Zhou, Weiwei, et al.
Pubblicazione: (2025)
di: Zhou, Weiwei, et al.
Pubblicazione: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
di: Yang, Kaicheng, et al.
Pubblicazione: (2024)
di: Yang, Kaicheng, et al.
Pubblicazione: (2024)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
di: Yang, Min, et al.
Pubblicazione: (2025)
di: Yang, Min, et al.
Pubblicazione: (2025)
CLIP-KD: An Empirical Study of CLIP Model Distillation
di: Yang, Chuanguang, et al.
Pubblicazione: (2023)
di: Yang, Chuanguang, et al.
Pubblicazione: (2023)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
di: Liu, Huimin, et al.
Pubblicazione: (2025)
di: Liu, Huimin, et al.
Pubblicazione: (2025)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
di: Huang, Weiquan, et al.
Pubblicazione: (2024)
di: Huang, Weiquan, et al.
Pubblicazione: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
di: Zhang, Beichen, et al.
Pubblicazione: (2024)
di: Zhang, Beichen, et al.
Pubblicazione: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
di: Song, Yehun, et al.
Pubblicazione: (2025)
di: Song, Yehun, et al.
Pubblicazione: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
di: Wang, Mengmeng, et al.
Pubblicazione: (2024)
di: Wang, Mengmeng, et al.
Pubblicazione: (2024)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
di: Du, Yao, et al.
Pubblicazione: (2025)
di: Du, Yao, et al.
Pubblicazione: (2025)
CLIP model is an Efficient Online Lifelong Learner
di: Wang, Leyuan, et al.
Pubblicazione: (2024)
di: Wang, Leyuan, et al.
Pubblicazione: (2024)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
Incremental Object Detection with CLIP
di: Huang, Ziyue, et al.
Pubblicazione: (2023)
di: Huang, Ziyue, et al.
Pubblicazione: (2023)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
di: Chaudhuri, Soumyabrata, et al.
Pubblicazione: (2024)
di: Chaudhuri, Soumyabrata, et al.
Pubblicazione: (2024)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
di: Ning, Shan, et al.
Pubblicazione: (2026)
di: Ning, Shan, et al.
Pubblicazione: (2026)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
di: Wang, Juan, et al.
Pubblicazione: (2026)
di: Wang, Juan, et al.
Pubblicazione: (2026)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
di: Lin, Kun-Yu, et al.
Pubblicazione: (2024)
di: Lin, Kun-Yu, et al.
Pubblicazione: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
di: Wang, Jingyun, et al.
Pubblicazione: (2024)
di: Wang, Jingyun, et al.
Pubblicazione: (2024)
Robust Light-Weight Facial Affective Behavior Recognition with CLIP
di: Lin, Li, et al.
Pubblicazione: (2024)
di: Lin, Li, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
di: Huang, Binhua, et al.
Pubblicazione: (2025) -
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
di: Maldonado, Gabriel, et al.
Pubblicazione: (2025) -
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
di: Huang, Binhua, et al.
Pubblicazione: (2025) -
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
di: Yao, Wendong, et al.
Pubblicazione: (2025) -
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
di: Yao, Wendong, et al.
Pubblicazione: (2025)