Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Verma, Dhruv, Roy, Debaditya, Fernando, Basura |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Predicting the Next Action by Modeling the Abstract Goal
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
por: Rajendiran, Ramanathan, et al.
Publicado: (2025)
por: Rajendiran, Ramanathan, et al.
Publicado: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
por: Rajendiran, Ramanathan, et al.
Publicado: (2023)
por: Rajendiran, Ramanathan, et al.
Publicado: (2023)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
por: Jaiswal, Shantanu, et al.
Publicado: (2024)
por: Jaiswal, Shantanu, et al.
Publicado: (2024)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
por: Ee, Yeo Keat, et al.
Publicado: (2026)
por: Ee, Yeo Keat, et al.
Publicado: (2026)
Situational Scene Graph for Structured Human-centric Situation Understanding
por: Sugandhika, Chinthani, et al.
Publicado: (2024)
por: Sugandhika, Chinthani, et al.
Publicado: (2024)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
por: Sugandhika, Chinthani, et al.
Publicado: (2025)
por: Sugandhika, Chinthani, et al.
Publicado: (2025)
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
por: Bangde, Yashwant Pravinrao, et al.
Publicado: (2026)
por: Bangde, Yashwant Pravinrao, et al.
Publicado: (2026)
Text-to-Image Generation Via Energy-Based CLIP
por: Ganz, Roy, et al.
Publicado: (2024)
por: Ganz, Roy, et al.
Publicado: (2024)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
por: Sugandhika, Chinthani, et al.
Publicado: (2025)
por: Sugandhika, Chinthani, et al.
Publicado: (2025)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
por: Li, Chen, et al.
Publicado: (2025)
por: Li, Chen, et al.
Publicado: (2025)
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
por: Zhang, Hao, et al.
Publicado: (2023)
por: Zhang, Hao, et al.
Publicado: (2023)
GLaRE: A Graph-based Landmark Region Embedding Network for Emotion Recognition
por: Maji, Debasis, et al.
Publicado: (2025)
por: Maji, Debasis, et al.
Publicado: (2025)
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
por: Ghildyal, Abhijay, et al.
Publicado: (2025)
por: Ghildyal, Abhijay, et al.
Publicado: (2025)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
por: Wei, Zhixiang, et al.
Publicado: (2025)
por: Wei, Zhixiang, et al.
Publicado: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
por: Fernando, Basura, et al.
Publicado: (2025)
por: Fernando, Basura, et al.
Publicado: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
por: Rawal, Ishaan Singh, et al.
Publicado: (2023)
por: Rawal, Ishaan Singh, et al.
Publicado: (2023)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
por: Parmar, Paritosh, et al.
Publicado: (2025)
por: Parmar, Paritosh, et al.
Publicado: (2025)
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
por: Tang, Zhenchen, et al.
Publicado: (2024)
por: Tang, Zhenchen, et al.
Publicado: (2024)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
por: Keat, Ee Yeo, et al.
Publicado: (2024)
por: Keat, Ee Yeo, et al.
Publicado: (2024)
Inferring Past Human Actions in Homes with Abductive Reasoning
por: Tan, Clement, et al.
Publicado: (2022)
por: Tan, Clement, et al.
Publicado: (2022)
Generating Key Postures of Bharatanatyam Adavus with Pose Estimation
por: Kamble, Jagadish Kashinath, et al.
Publicado: (2026)
por: Kamble, Jagadish Kashinath, et al.
Publicado: (2026)
Leveraging Self-Supervised Features for Efficient Flooded Region Identification in UAV Aerial Images
por: Deb, Dibyabha, et al.
Publicado: (2025)
por: Deb, Dibyabha, et al.
Publicado: (2025)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
por: Song, Yehun, et al.
Publicado: (2025)
por: Song, Yehun, et al.
Publicado: (2025)
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
por: Benschop, Pascal, et al.
Publicado: (2026)
por: Benschop, Pascal, et al.
Publicado: (2026)
MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video
por: Guo, Xiaoqing, et al.
Publicado: (2024)
por: Guo, Xiaoqing, et al.
Publicado: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
por: Zhu, Wencheng, et al.
Publicado: (2025)
por: Zhu, Wencheng, et al.
Publicado: (2025)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
por: Zhu, Wenqi, et al.
Publicado: (2024)
por: Zhu, Wenqi, et al.
Publicado: (2024)
CLIP-IT: CLIP-based Pairing for Histology Images Classification
por: Karimian, Banafsheh, et al.
Publicado: (2025)
por: Karimian, Banafsheh, et al.
Publicado: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
por: Chaffin, Antoine, et al.
Publicado: (2024)
por: Chaffin, Antoine, et al.
Publicado: (2024)
Context-Aware Pesticide Recommendation via Few-Shot Pest Recognition for Precision Agriculture
por: Ghosh, Anirudha, et al.
Publicado: (2026)
por: Ghosh, Anirudha, et al.
Publicado: (2026)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
por: Huang, Binhua, et al.
Publicado: (2025)
por: Huang, Binhua, et al.
Publicado: (2025)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
por: Zhang, Zilun, et al.
Publicado: (2022)
por: Zhang, Zilun, et al.
Publicado: (2022)
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
por: Agarwal, Lakshita, et al.
Publicado: (2025)
por: Agarwal, Lakshita, et al.
Publicado: (2025)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
por: Singh, Darshan, et al.
Publicado: (2024)
por: Singh, Darshan, et al.
Publicado: (2024)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
por: Biswas, Shristi Das, et al.
Publicado: (2025)
por: Biswas, Shristi Das, et al.
Publicado: (2025)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
por: Wang, Juan, et al.
Publicado: (2026)
por: Wang, Juan, et al.
Publicado: (2026)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
por: Lv, Song-Lin, et al.
Publicado: (2025)
por: Lv, Song-Lin, et al.
Publicado: (2025)
Ejemplares similares
-
Predicting the Next Action by Modeling the Abstract Goal
por: Roy, Debaditya, et al.
Publicado: (2022) -
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
por: Rajendiran, Ramanathan, et al.
Publicado: (2025) -
Interaction Region Visual Transformer for Egocentric Action Anticipation
por: Roy, Debaditya, et al.
Publicado: (2022) -
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
por: Rajendiran, Ramanathan, et al.
Publicado: (2023) -
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
por: Jaiswal, Shantanu, et al.
Publicado: (2024)