AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Basappa, Aahana, Goel, Pranay, Karra, Anusri, Karra, Anish, Gilmore, Asa, Zhu, Kevin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
di: Zhang, Junyang, et al.
Pubblicazione: (2025)
di: Zhang, Junyang, et al.
Pubblicazione: (2025)
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
di: Erol, Mehmet Kaan
Pubblicazione: (2026)
di: Erol, Mehmet Kaan
Pubblicazione: (2026)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
di: Singh, Anshul, et al.
Pubblicazione: (2025)
di: Singh, Anshul, et al.
Pubblicazione: (2025)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
di: Mak, Chak-Wing, et al.
Pubblicazione: (2026)
di: Mak, Chak-Wing, et al.
Pubblicazione: (2026)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
di: Azad, Asif, et al.
Pubblicazione: (2025)
di: Azad, Asif, et al.
Pubblicazione: (2025)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
di: Jian, Ai, et al.
Pubblicazione: (2025)
di: Jian, Ai, et al.
Pubblicazione: (2025)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
di: Popordanoska, Teodora, et al.
Pubblicazione: (2025)
di: Popordanoska, Teodora, et al.
Pubblicazione: (2025)
CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems
di: Tian, Yonglin, et al.
Pubblicazione: (2026)
di: Tian, Yonglin, et al.
Pubblicazione: (2026)
Discovering Failure Modes in Vision-Language Models using RL
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
di: Tian, Yingjie, et al.
Pubblicazione: (2023)
di: Tian, Yingjie, et al.
Pubblicazione: (2023)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
di: Tang, Zecong, et al.
Pubblicazione: (2026)
di: Tang, Zecong, et al.
Pubblicazione: (2026)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
di: Ballout, Mohamad, et al.
Pubblicazione: (2025)
di: Ballout, Mohamad, et al.
Pubblicazione: (2025)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
di: Singha, Mainak, et al.
Pubblicazione: (2026)
di: Singha, Mainak, et al.
Pubblicazione: (2026)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
di: Mayer, Julius, et al.
Pubblicazione: (2025)
di: Mayer, Julius, et al.
Pubblicazione: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
di: Faraz, Ali, et al.
Pubblicazione: (2025)
di: Faraz, Ali, et al.
Pubblicazione: (2025)
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
di: Huang, Hai, et al.
Pubblicazione: (2024)
di: Huang, Hai, et al.
Pubblicazione: (2024)
A Study of Failure Modes in Two-Stage Human-Object Interaction Detection
di: Wang, Lemeng, et al.
Pubblicazione: (2026)
di: Wang, Lemeng, et al.
Pubblicazione: (2026)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
Source-Free Cross-Modal Knowledge Transfer by Unleashing the Potential of Task-Irrelevant Data
di: Zhu, Jinjing, et al.
Pubblicazione: (2024)
di: Zhu, Jinjing, et al.
Pubblicazione: (2024)
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
di: Berian, Alex, et al.
Pubblicazione: (2025)
di: Berian, Alex, et al.
Pubblicazione: (2025)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
di: Xia, Shao-Jun, et al.
Pubblicazione: (2025)
di: Xia, Shao-Jun, et al.
Pubblicazione: (2025)
Caption This, Reason That: VLMs Caught in the Middle
di: Weng, Zihan, et al.
Pubblicazione: (2025)
di: Weng, Zihan, et al.
Pubblicazione: (2025)
CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
di: Zhang, Dongxu, et al.
Pubblicazione: (2026)
di: Zhang, Dongxu, et al.
Pubblicazione: (2026)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
di: Bhaskar, Paramananda, et al.
Pubblicazione: (2026)
di: Bhaskar, Paramananda, et al.
Pubblicazione: (2026)
SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis
di: Pramanik, Rishav, et al.
Pubblicazione: (2026)
di: Pramanik, Rishav, et al.
Pubblicazione: (2026)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
di: Lian, Weitong, et al.
Pubblicazione: (2026)
di: Lian, Weitong, et al.
Pubblicazione: (2026)
Cross-Modal Learning of Housing Quality in Amsterdam
di: Levering, Alex, et al.
Pubblicazione: (2024)
di: Levering, Alex, et al.
Pubblicazione: (2024)
Evaluating Compositional Generalisation in VLMs and Diffusion Models
di: Pearson, Beth, et al.
Pubblicazione: (2025)
di: Pearson, Beth, et al.
Pubblicazione: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
di: Kim, Jeonghyeon, et al.
Pubblicazione: (2025)
di: Kim, Jeonghyeon, et al.
Pubblicazione: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
di: Zhou, Hefeng, et al.
Pubblicazione: (2026)
di: Zhou, Hefeng, et al.
Pubblicazione: (2026)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
di: Yan, Hanqi, et al.
Pubblicazione: (2025)
di: Yan, Hanqi, et al.
Pubblicazione: (2025)
SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability
di: Nasiri-Sarvi, Ali, et al.
Pubblicazione: (2025)
di: Nasiri-Sarvi, Ali, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
di: Zhang, Junyang, et al.
Pubblicazione: (2025) -
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
di: Erol, Mehmet Kaan
Pubblicazione: (2026) -
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
di: Singh, Anshul, et al.
Pubblicazione: (2025) -
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
di: Hong, Hyesoo, et al.
Pubblicazione: (2026) -
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
di: Mak, Chak-Wing, et al.
Pubblicazione: (2026)