Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gaur, Manu, S, Darshan Singh, Tapaswi, Makarand |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
von: Gaur, Manu, et al.
Veröffentlicht: (2024)
von: Gaur, Manu, et al.
Veröffentlicht: (2024)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
von: Singh, Darshan, et al.
Veröffentlicht: (2024)
von: Singh, Darshan, et al.
Veröffentlicht: (2024)
Steerable Visual Representations
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
"Previously on ..." From Recaps to Story Summarization
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
von: Kala, Divy, et al.
Veröffentlicht: (2025)
von: Kala, Divy, et al.
Veröffentlicht: (2025)
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
von: Darur, Balaji, et al.
Veröffentlicht: (2026)
von: Darur, Balaji, et al.
Veröffentlicht: (2026)
MALeR: Improving Compositional Fidelity in Layout-Guided Generation
von: Saxena, Shivank, et al.
Veröffentlicht: (2025)
von: Saxena, Shivank, et al.
Veröffentlicht: (2025)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
von: Saravanan, Darshana, et al.
Veröffentlicht: (2024)
von: Saravanan, Darshana, et al.
Veröffentlicht: (2024)
Investigating Mechanisms for In-Context Vision Language Binding
von: Saravanan, Darshana, et al.
Veröffentlicht: (2025)
von: Saravanan, Darshana, et al.
Veröffentlicht: (2025)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
von: Kumar, Prajneya, et al.
Veröffentlicht: (2023)
von: Kumar, Prajneya, et al.
Veröffentlicht: (2023)
MICap: A Unified Model for Identity-aware Movie Descriptions
von: Raajesh, Haran, et al.
Veröffentlicht: (2024)
von: Raajesh, Haran, et al.
Veröffentlicht: (2024)
STRinGS: Selective Text Refinement in Gaussian Splatting
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
von: Bagad, Piyush, et al.
Veröffentlicht: (2024)
von: Bagad, Piyush, et al.
Veröffentlicht: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
Autonomous Crack Detection using Deep Learning on Synthetic Thermogram Datasets
von: Pimpalkhare, Chinmay Makarand, et al.
Veröffentlicht: (2024)
von: Pimpalkhare, Chinmay Makarand, et al.
Veröffentlicht: (2024)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
von: Wei, Lai, et al.
Veröffentlicht: (2024)
von: Wei, Lai, et al.
Veröffentlicht: (2024)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
von: Zou, Yueying, et al.
Veröffentlicht: (2025)
von: Zou, Yueying, et al.
Veröffentlicht: (2025)
Beyond Normal References: Discriminative Few-Shot Anomaly Detection
von: Wang, Huan, et al.
Veröffentlicht: (2026)
von: Wang, Huan, et al.
Veröffentlicht: (2026)
Described Spatial-Temporal Video Detection
von: Ji, Wei, et al.
Veröffentlicht: (2024)
von: Ji, Wei, et al.
Veröffentlicht: (2024)
Synergizing Discriminative Exemplars and Self-Refined Experience for MLLM-based In-Context Learning in Medical Diagnosis
von: Zhao, Wenkai, et al.
Veröffentlicht: (2026)
von: Zhao, Wenkai, et al.
Veröffentlicht: (2026)
Abnormal Event Detection In Videos Using Deep Embedding
von: Venkatrayappa, Darshan
Veröffentlicht: (2024)
von: Venkatrayappa, Darshan
Veröffentlicht: (2024)
Described Object Detection: Liberating Object Detection with Flexible Expressions
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
HLTCOE Evaluation Team at TREC 2025: VQA Track
von: Zhang, Dengjia, et al.
Veröffentlicht: (2025)
von: Zhang, Dengjia, et al.
Veröffentlicht: (2025)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection
von: Zhang, Guoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Guoyi, et al.
Veröffentlicht: (2025)
DescribeEarth: Describe Anything for Remote Sensing Images
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
Moment and Highlight Detection via MLLM Frame Segmentation
von: Jiwanta, I Putu Andika Bagas, et al.
Veröffentlicht: (2025)
von: Jiwanta, I Putu Andika Bagas, et al.
Veröffentlicht: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
von: Zhao, Zaiying, et al.
Veröffentlicht: (2025)
von: Zhao, Zaiying, et al.
Veröffentlicht: (2025)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
On the Role of Visual Grounding in VQA
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
von: Reich, Daniel, et al.
Veröffentlicht: (2024)
VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection
von: Han, Hui, et al.
Veröffentlicht: (2026)
von: Han, Hui, et al.
Veröffentlicht: (2026)
ChangeMinds: Multi-task Framework for Detecting and Describing Changes in Remote Sensing
von: Wang, Yuduo, et al.
Veröffentlicht: (2024)
von: Wang, Yuduo, et al.
Veröffentlicht: (2024)
Describe Anything in Medical Images
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
von: Wen, Haiquan, et al.
Veröffentlicht: (2025)
von: Wen, Haiquan, et al.
Veröffentlicht: (2025)
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
von: Gaur, Manu, et al.
Veröffentlicht: (2024) -
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
von: Singh, Darshan, et al.
Veröffentlicht: (2024) -
Steerable Visual Representations
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026) -
"Previously on ..." From Recaps to Story Summarization
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2024) -
What You See is What You Ask: Evaluating Audio Descriptions
von: Kala, Divy, et al.
Veröffentlicht: (2025)