SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Darshan, Khan, Zeeshan, Tapaswi, Makarand |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
por: Gaur, Manu, et al.
Publicado: (2024)
por: Gaur, Manu, et al.
Publicado: (2024)
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
por: Gaur, Manu, et al.
Publicado: (2024)
por: Gaur, Manu, et al.
Publicado: (2024)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
por: Saravanan, Darshana, et al.
Publicado: (2024)
por: Saravanan, Darshana, et al.
Publicado: (2024)
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
por: Darur, Balaji, et al.
Publicado: (2026)
por: Darur, Balaji, et al.
Publicado: (2026)
MICap: A Unified Model for Identity-aware Movie Descriptions
por: Raajesh, Haran, et al.
Publicado: (2024)
por: Raajesh, Haran, et al.
Publicado: (2024)
"Previously on ..." From Recaps to Story Summarization
por: Singh, Aditya Kumar, et al.
Publicado: (2024)
por: Singh, Aditya Kumar, et al.
Publicado: (2024)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
por: Kumar, Prajneya, et al.
Publicado: (2023)
por: Kumar, Prajneya, et al.
Publicado: (2023)
MALeR: Improving Compositional Fidelity in Layout-Guided Generation
por: Saxena, Shivank, et al.
Publicado: (2025)
por: Saxena, Shivank, et al.
Publicado: (2025)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
por: Alyami, Sarah, et al.
Publicado: (2025)
por: Alyami, Sarah, et al.
Publicado: (2025)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
por: Zhu, Wencheng, et al.
Publicado: (2025)
por: Zhu, Wencheng, et al.
Publicado: (2025)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
por: Du, Yao, et al.
Publicado: (2025)
por: Du, Yao, et al.
Publicado: (2025)
Investigating Mechanisms for In-Context Vision Language Binding
por: Saravanan, Darshana, et al.
Publicado: (2025)
por: Saravanan, Darshana, et al.
Publicado: (2025)
Label-Efficient Chest X-ray Diagnosis via Partial CLIP Adaptation
por: Dalsania, Heet Nitinkumar
Publicado: (2025)
por: Dalsania, Heet Nitinkumar
Publicado: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
por: Zhang, Kangjie, et al.
Publicado: (2026)
por: Zhang, Kangjie, et al.
Publicado: (2026)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
por: Huang, Binhua, et al.
Publicado: (2025)
por: Huang, Binhua, et al.
Publicado: (2025)
AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
por: Fang, Qingqing, et al.
Publicado: (2025)
por: Fang, Qingqing, et al.
Publicado: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
por: Kala, Divy, et al.
Publicado: (2025)
por: Kala, Divy, et al.
Publicado: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
por: Yang, Kaicheng, et al.
Publicado: (2024)
por: Yang, Kaicheng, et al.
Publicado: (2024)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
por: Silva, Sathira, et al.
Publicado: (2025)
por: Silva, Sathira, et al.
Publicado: (2025)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
por: Nawaz, Umair, et al.
Publicado: (2024)
por: Nawaz, Umair, et al.
Publicado: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
por: Wang, Jingyun, et al.
Publicado: (2024)
por: Wang, Jingyun, et al.
Publicado: (2024)
Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
por: Ali, Eman, et al.
Publicado: (2025)
por: Ali, Eman, et al.
Publicado: (2025)
CLIPCleaner: Cleaning Noisy Labels with CLIP
por: Feng, Chen, et al.
Publicado: (2024)
por: Feng, Chen, et al.
Publicado: (2024)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
por: Ali, Muhammad, et al.
Publicado: (2024)
por: Ali, Muhammad, et al.
Publicado: (2024)
DeCLIP: Decoupled Prompting for CLIP-based Multi-Label Class-Incremental Learning
por: Du, Kaile, et al.
Publicado: (2025)
por: Du, Kaile, et al.
Publicado: (2025)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
por: Zhu, Wenqi, et al.
Publicado: (2024)
por: Zhu, Wenqi, et al.
Publicado: (2024)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
por: Chung, Jeannie, et al.
Publicado: (2026)
por: Chung, Jeannie, et al.
Publicado: (2026)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
por: Zhang, Dengke, et al.
Publicado: (2024)
por: Zhang, Dengke, et al.
Publicado: (2024)
CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification
por: Lu, Liping, et al.
Publicado: (2025)
por: Lu, Liping, et al.
Publicado: (2025)
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning
por: Menghini, Cristina, et al.
Publicado: (2023)
por: Menghini, Cristina, et al.
Publicado: (2023)
CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values
por: Koleilat, Taha, et al.
Publicado: (2025)
por: Koleilat, Taha, et al.
Publicado: (2025)
HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA
por: Dibitonto, Francesco, et al.
Publicado: (2026)
por: Dibitonto, Francesco, et al.
Publicado: (2026)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
por: Li, Yinqi, et al.
Publicado: (2025)
por: Li, Yinqi, et al.
Publicado: (2025)
CLIP Adaptation by Intra-modal Overlap Reduction
por: Kravets, Alexey, et al.
Publicado: (2024)
por: Kravets, Alexey, et al.
Publicado: (2024)
Rethinking Domain Adaptation and Generalization in the Era of CLIP
por: Feng, Ruoyu, et al.
Publicado: (2024)
por: Feng, Ruoyu, et al.
Publicado: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
por: Wang, Jiapeng, et al.
Publicado: (2024)
por: Wang, Jiapeng, et al.
Publicado: (2024)
DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
por: Yang, Zhiwei, et al.
Publicado: (2026)
por: Yang, Zhiwei, et al.
Publicado: (2026)
CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation
por: Jalilian, Shayan, et al.
Publicado: (2026)
por: Jalilian, Shayan, et al.
Publicado: (2026)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
por: Jon, Hyo Jin, et al.
Publicado: (2026)
por: Jon, Hyo Jin, et al.
Publicado: (2026)
CLIP-driven Zero-shot Learning with Ambiguous Labels
por: Fan, Jinfu, et al.
Publicado: (2026)
por: Fan, Jinfu, et al.
Publicado: (2026)
Ejemplares similares
-
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
por: Gaur, Manu, et al.
Publicado: (2024) -
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
por: Gaur, Manu, et al.
Publicado: (2024) -
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
por: Saravanan, Darshana, et al.
Publicado: (2024) -
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
por: Darur, Balaji, et al.
Publicado: (2026) -
MICap: A Unified Model for Identity-aware Movie Descriptions
por: Raajesh, Haran, et al.
Publicado: (2024)