MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Binhua, Wang, Ni, Pakrashi, Arjun, Dev, Soumyabrata |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
por: Huang, Binhua, et al.
Publicado: (2025)
por: Huang, Binhua, et al.
Publicado: (2025)
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
por: Maldonado, Gabriel, et al.
Publicado: (2025)
por: Maldonado, Gabriel, et al.
Publicado: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
por: Huang, Binhua, et al.
Publicado: (2025)
por: Huang, Binhua, et al.
Publicado: (2025)
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
por: Yao, Wendong, et al.
Publicado: (2025)
por: Yao, Wendong, et al.
Publicado: (2025)
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
por: Yao, Wendong, et al.
Publicado: (2025)
por: Yao, Wendong, et al.
Publicado: (2025)
LiteEmbed: Adapting CLIP to Rare Classes
por: Agarwal, Aishwarya, et al.
Publicado: (2026)
por: Agarwal, Aishwarya, et al.
Publicado: (2026)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
por: Alyami, Sarah, et al.
Publicado: (2025)
por: Alyami, Sarah, et al.
Publicado: (2025)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
por: Wang, Jiapeng, et al.
Publicado: (2024)
por: Wang, Jiapeng, et al.
Publicado: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
por: Ahmad, Shahzad, et al.
Publicado: (2023)
por: Ahmad, Shahzad, et al.
Publicado: (2023)
A Deep Learning Approach for Spatio-Temporal Forecasting of InSAR Ground Deformation in Eastern Ireland
por: Yao, Wendong, et al.
Publicado: (2025)
por: Yao, Wendong, et al.
Publicado: (2025)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
por: Liu, Mushui, et al.
Publicado: (2024)
por: Liu, Mushui, et al.
Publicado: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
por: Zhu, Wencheng, et al.
Publicado: (2025)
por: Zhu, Wencheng, et al.
Publicado: (2025)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
por: Singh, Darshan, et al.
Publicado: (2024)
por: Singh, Darshan, et al.
Publicado: (2024)
AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
por: Zhong, Enmin, et al.
Publicado: (2025)
por: Zhong, Enmin, et al.
Publicado: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
por: Zhao, Weiheng, et al.
Publicado: (2025)
por: Zhao, Weiheng, et al.
Publicado: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
por: Zhang, Kangjie, et al.
Publicado: (2026)
por: Zhang, Kangjie, et al.
Publicado: (2026)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
por: Song, Dan, et al.
Publicado: (2023)
por: Song, Dan, et al.
Publicado: (2023)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
por: Zhu, Wenqi, et al.
Publicado: (2024)
por: Zhu, Wenqi, et al.
Publicado: (2024)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
por: Chung, Jeannie, et al.
Publicado: (2026)
por: Chung, Jeannie, et al.
Publicado: (2026)
Emotion Recognition with CLIP and Sequential Learning
por: Zhou, Weiwei, et al.
Publicado: (2025)
por: Zhou, Weiwei, et al.
Publicado: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
por: Yang, Kaicheng, et al.
Publicado: (2024)
por: Yang, Kaicheng, et al.
Publicado: (2024)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
por: Yang, Min, et al.
Publicado: (2025)
por: Yang, Min, et al.
Publicado: (2025)
CLIP-KD: An Empirical Study of CLIP Model Distillation
por: Yang, Chuanguang, et al.
Publicado: (2023)
por: Yang, Chuanguang, et al.
Publicado: (2023)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
por: Liu, Huimin, et al.
Publicado: (2025)
por: Liu, Huimin, et al.
Publicado: (2025)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
por: Huang, Weiquan, et al.
Publicado: (2024)
por: Huang, Weiquan, et al.
Publicado: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
por: Zhang, Beichen, et al.
Publicado: (2024)
por: Zhang, Beichen, et al.
Publicado: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
por: Song, Yehun, et al.
Publicado: (2025)
por: Song, Yehun, et al.
Publicado: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
por: Wang, Mengmeng, et al.
Publicado: (2024)
por: Wang, Mengmeng, et al.
Publicado: (2024)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
por: Du, Yao, et al.
Publicado: (2025)
por: Du, Yao, et al.
Publicado: (2025)
CLIP model is an Efficient Online Lifelong Learner
por: Wang, Leyuan, et al.
Publicado: (2024)
por: Wang, Leyuan, et al.
Publicado: (2024)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
por: Jon, Hyo Jin, et al.
Publicado: (2026)
por: Jon, Hyo Jin, et al.
Publicado: (2026)
Incremental Object Detection with CLIP
por: Huang, Ziyue, et al.
Publicado: (2023)
por: Huang, Ziyue, et al.
Publicado: (2023)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
por: Zhang, Jihai, et al.
Publicado: (2024)
por: Zhang, Jihai, et al.
Publicado: (2024)
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
por: Chaudhuri, Soumyabrata, et al.
Publicado: (2024)
por: Chaudhuri, Soumyabrata, et al.
Publicado: (2024)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
por: Ning, Shan, et al.
Publicado: (2026)
por: Ning, Shan, et al.
Publicado: (2026)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
por: Wang, Juan, et al.
Publicado: (2026)
por: Wang, Juan, et al.
Publicado: (2026)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
por: Lin, Kun-Yu, et al.
Publicado: (2024)
por: Lin, Kun-Yu, et al.
Publicado: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
por: Wang, Jingyun, et al.
Publicado: (2024)
por: Wang, Jingyun, et al.
Publicado: (2024)
Robust Light-Weight Facial Affective Behavior Recognition with CLIP
por: Lin, Li, et al.
Publicado: (2024)
por: Lin, Li, et al.
Publicado: (2024)
Ejemplares similares
-
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
por: Huang, Binhua, et al.
Publicado: (2025) -
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
por: Maldonado, Gabriel, et al.
Publicado: (2025) -
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
por: Huang, Binhua, et al.
Publicado: (2025) -
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
por: Yao, Wendong, et al.
Publicado: (2025) -
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
por: Yao, Wendong, et al.
Publicado: (2025)