VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saxena, Rohit, Suglia, Alessandro, Minervini, Pasquale |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
Same Answer, Different Representations: Hidden instability in VLMs
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
von: Lorenz, Peter, et al.
Veröffentlicht: (2021)
von: Lorenz, Peter, et al.
Veröffentlicht: (2021)
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
von: Wang, Sibo, et al.
Veröffentlicht: (2024)
von: Wang, Sibo, et al.
Veröffentlicht: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
von: Suglia, Alessandro, et al.
Veröffentlicht: (2024)
von: Suglia, Alessandro, et al.
Veröffentlicht: (2024)
RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
von: Park, Junghyun, et al.
Veröffentlicht: (2025)
von: Park, Junghyun, et al.
Veröffentlicht: (2025)
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions
von: Ma, Sihan, et al.
Veröffentlicht: (2024)
von: Ma, Sihan, et al.
Veröffentlicht: (2024)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
von: Meng, Zijie, et al.
Veröffentlicht: (2025)
von: Meng, Zijie, et al.
Veröffentlicht: (2025)
NIC-RobustBench: A Comprehensive Open-Source Toolkit for Neural Image Compression and Robustness Analysis
von: Bychkov, Georgii, et al.
Veröffentlicht: (2025)
von: Bychkov, Georgii, et al.
Veröffentlicht: (2025)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
von: Cai, Jie, et al.
Veröffentlicht: (2025)
von: Cai, Jie, et al.
Veröffentlicht: (2025)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
von: Chen, Ce, et al.
Veröffentlicht: (2026)
von: Chen, Ce, et al.
Veröffentlicht: (2026)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
von: Bao, Han, et al.
Veröffentlicht: (2024)
von: Bao, Han, et al.
Veröffentlicht: (2024)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
von: Luo, Zhiming, et al.
Veröffentlicht: (2026)
von: Luo, Zhiming, et al.
Veröffentlicht: (2026)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
RadVLM: A Multitask Conversational Vision-Language Model for Radiology
von: Deperrois, Nicolas, et al.
Veröffentlicht: (2025)
von: Deperrois, Nicolas, et al.
Veröffentlicht: (2025)
VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
von: Lin, Zihang, et al.
Veröffentlicht: (2026)
von: Lin, Zihang, et al.
Veröffentlicht: (2026)
NeuroVLM-Bench: Evaluation of Vision-Enabled Large Language Models for Clinical Reasoning in Neurological Disorders
von: Dineva, Katarina Trojachanec, et al.
Veröffentlicht: (2026)
von: Dineva, Katarina Trojachanec, et al.
Veröffentlicht: (2026)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
von: Mathew, Athul M., et al.
Veröffentlicht: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
von: Saxena, Rohit, et al.
Veröffentlicht: (2025) -
DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation
von: Wu, Yifan, et al.
Veröffentlicht: (2024) -
Same Answer, Different Representations: Hidden instability in VLMs
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026) -
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
von: Lorenz, Peter, et al.
Veröffentlicht: (2021)