MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Binhua, Wang, Ni, Pakrashi, Arjun, Dev, Soumyabrata |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
LiteEmbed: Adapting CLIP to Rare Classes
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2026)
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2026)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
A Deep Learning Approach for Spatio-Temporal Forecasting of InSAR Ground Deformation in Eastern Ireland
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
von: Yao, Wendong, et al.
Veröffentlicht: (2025)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
von: Singh, Darshan, et al.
Veröffentlicht: (2024)
von: Singh, Darshan, et al.
Veröffentlicht: (2024)
AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
von: Zhong, Enmin, et al.
Veröffentlicht: (2025)
von: Zhong, Enmin, et al.
Veröffentlicht: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
von: Song, Dan, et al.
Veröffentlicht: (2023)
von: Song, Dan, et al.
Veröffentlicht: (2023)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
von: Zhu, Wenqi, et al.
Veröffentlicht: (2024)
von: Zhu, Wenqi, et al.
Veröffentlicht: (2024)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
von: Chung, Jeannie, et al.
Veröffentlicht: (2026)
von: Chung, Jeannie, et al.
Veröffentlicht: (2026)
Emotion Recognition with CLIP and Sequential Learning
von: Zhou, Weiwei, et al.
Veröffentlicht: (2025)
von: Zhou, Weiwei, et al.
Veröffentlicht: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
von: Yang, Min, et al.
Veröffentlicht: (2025)
von: Yang, Min, et al.
Veröffentlicht: (2025)
CLIP-KD: An Empirical Study of CLIP Model Distillation
von: Yang, Chuanguang, et al.
Veröffentlicht: (2023)
von: Yang, Chuanguang, et al.
Veröffentlicht: (2023)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
von: Liu, Huimin, et al.
Veröffentlicht: (2025)
von: Liu, Huimin, et al.
Veröffentlicht: (2025)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
von: Song, Yehun, et al.
Veröffentlicht: (2025)
von: Song, Yehun, et al.
Veröffentlicht: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
von: Wang, Mengmeng, et al.
Veröffentlicht: (2024)
von: Wang, Mengmeng, et al.
Veröffentlicht: (2024)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
von: Du, Yao, et al.
Veröffentlicht: (2025)
von: Du, Yao, et al.
Veröffentlicht: (2025)
CLIP model is an Efficient Online Lifelong Learner
von: Wang, Leyuan, et al.
Veröffentlicht: (2024)
von: Wang, Leyuan, et al.
Veröffentlicht: (2024)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2026)
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2026)
Incremental Object Detection with CLIP
von: Huang, Ziyue, et al.
Veröffentlicht: (2023)
von: Huang, Ziyue, et al.
Veröffentlicht: (2023)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
von: Chaudhuri, Soumyabrata, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Soumyabrata, et al.
Veröffentlicht: (2024)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
von: Ning, Shan, et al.
Veröffentlicht: (2026)
von: Ning, Shan, et al.
Veröffentlicht: (2026)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
von: Wang, Juan, et al.
Veröffentlicht: (2026)
von: Wang, Juan, et al.
Veröffentlicht: (2026)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
Robust Light-Weight Facial Affective Behavior Recognition with CLIP
von: Lin, Li, et al.
Veröffentlicht: (2024)
von: Lin, Li, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
von: Huang, Binhua, et al.
Veröffentlicht: (2025) -
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025) -
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
von: Huang, Binhua, et al.
Veröffentlicht: (2025) -
Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
von: Yao, Wendong, et al.
Veröffentlicht: (2025) -
A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe
von: Yao, Wendong, et al.
Veröffentlicht: (2025)