How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Khattak, Muhammad Uzair, Naeem, Muhammad Ferjad, Hassan, Jameel, Naseer, Muzammal, Tombari, Federico, Khan, Fahad Shahbaz, Khan, Salman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
by: Hassan, Jameel, et al.
Published: (2023)
by: Hassan, Jameel, et al.
Published: (2023)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
by: Maaz, Muhammad, et al.
Published: (2023)
by: Maaz, Muhammad, et al.
Published: (2023)
How Good is my Histopathology Vision-Language Foundation Model? A Holistic Benchmark
by: Majzoub, Roba Al, et al.
Published: (2025)
by: Majzoub, Roba Al, et al.
Published: (2025)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
Towards Evaluating the Robustness of Visual State Space Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2023)
by: Kunhimon, Shahina, et al.
Published: (2023)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
by: Hanif, Asif, et al.
Published: (2024)
by: Hanif, Asif, et al.
Published: (2024)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
by: Gani, Hanan, et al.
Published: (2024)
by: Gani, Hanan, et al.
Published: (2024)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
On Evaluating Adversarial Robustness of Volumetric Medical Segmentation Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
by: Nawaz, Umair, et al.
Published: (2024)
by: Nawaz, Umair, et al.
Published: (2024)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
Efficient Localized Adaptation of Neural Weather Forecasting: A Case Study in the MENA Region
by: Munir, Muhammad Akhtar, et al.
Published: (2024)
by: Munir, Muhammad Akhtar, et al.
Published: (2024)
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026)
by: Nawaz, Umair, et al.
Published: (2026)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
by: Munasinghe, Shehan, et al.
Published: (2024)
by: Munasinghe, Shehan, et al.
Published: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
by: Ghaboura, Sara, et al.
Published: (2025)
by: Ghaboura, Sara, et al.
Published: (2025)
GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
Efficient IoT Intrusion Detection with an Improved Attention-Based CNN-BiLSTM Architecture
by: Naeem, Amna, et al.
Published: (2025)
by: Naeem, Amna, et al.
Published: (2025)
Similar Items
-
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024) -
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024) -
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024) -
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
by: Hassan, Jameel, et al.
Published: (2023) -
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)