Composed Video Retrieval via Enriched Context and Discriminative Embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thawakar, Omkar, Naseer, Muzammal, Anwer, Rao Muhammad, Khan, Salman, Felsberg, Michael, Shah, Mubarak, Khan, Fahad Shahbaz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
von: Hanif, Asif, et al.
Veröffentlicht: (2024)
von: Hanif, Asif, et al.
Veröffentlicht: (2024)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
Language Guided Domain Generalized Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2024)
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2024)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
von: Nawaz, Umair, et al.
Veröffentlicht: (2024)
von: Nawaz, Umair, et al.
Veröffentlicht: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
CoVR-R:Reason-Aware Composed Video Retrieval
von: Thawakar, Omkar, et al.
Veröffentlicht: (2026)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2026)
Multi-modal Generation via Cross-Modal In-Context Learning
von: Kumar, Amandeep, et al.
Veröffentlicht: (2024)
von: Kumar, Amandeep, et al.
Veröffentlicht: (2024)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
von: Gani, Hanan, et al.
Veröffentlicht: (2024)
von: Gani, Hanan, et al.
Veröffentlicht: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
von: Wang, Tong, et al.
Veröffentlicht: (2025)
von: Wang, Tong, et al.
Veröffentlicht: (2025)
Towards Evaluating the Robustness of Visual State Space Models
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
von: Patle, Shubham, et al.
Veröffentlicht: (2026)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
von: Nawaz, Umair, et al.
Veröffentlicht: (2025)
von: Nawaz, Umair, et al.
Veröffentlicht: (2025)
AIN: The Arabic INclusive Large Multimodal Model
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
von: Cui, Yuning, et al.
Veröffentlicht: (2024)
von: Cui, Yuning, et al.
Veröffentlicht: (2024)
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025) -
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025) -
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024) -
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
von: Hanif, Asif, et al.
Veröffentlicht: (2024) -
CDChat: A Large Multimodal Model for Remote Sensing Change Description
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)