VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Hyeonsu, Bao, Emily, Goswami, Anjan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
di: Li, Frank, et al.
Pubblicazione: (2025)
di: Li, Frank, et al.
Pubblicazione: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2025)
di: Pani, Anupam, et al.
Pubblicazione: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
di: Wang, Shengao, et al.
Pubblicazione: (2025)
di: Wang, Shengao, et al.
Pubblicazione: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
di: Kang, Inha, et al.
Pubblicazione: (2025)
di: Kang, Inha, et al.
Pubblicazione: (2025)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
Evaluating Compositional Generalisation in VLMs and Diffusion Models
di: Pearson, Beth, et al.
Pubblicazione: (2025)
di: Pearson, Beth, et al.
Pubblicazione: (2025)
Replace-then-Perturb: Targeted Adversarial Attacks With Visual Reasoning for Vision-Language Models
di: Jang, Jonggyu, et al.
Pubblicazione: (2024)
di: Jang, Jonggyu, et al.
Pubblicazione: (2024)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
di: He, Zheqi, et al.
Pubblicazione: (2025)
di: He, Zheqi, et al.
Pubblicazione: (2025)
An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
di: Lee, Yuxiao, et al.
Pubblicazione: (2025)
di: Lee, Yuxiao, et al.
Pubblicazione: (2025)
Trust but Verify: Programmatic VLM Evaluation in the Wild
di: Prabhu, Viraj, et al.
Pubblicazione: (2024)
di: Prabhu, Viraj, et al.
Pubblicazione: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
di: Saxena, Rohit, et al.
Pubblicazione: (2026)
di: Saxena, Rohit, et al.
Pubblicazione: (2026)
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
di: Han, Shuhao, et al.
Pubblicazione: (2024)
di: Han, Shuhao, et al.
Pubblicazione: (2024)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
di: Jian, Ai, et al.
Pubblicazione: (2025)
di: Jian, Ai, et al.
Pubblicazione: (2025)
DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
di: Meng, Zijie, et al.
Pubblicazione: (2025)
di: Meng, Zijie, et al.
Pubblicazione: (2025)
Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction
di: Bejerano, Emily, et al.
Pubblicazione: (2026)
di: Bejerano, Emily, et al.
Pubblicazione: (2026)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)
di: Kamath, Amita, et al.
Pubblicazione: (2025)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
di: Peng, Tianhao, et al.
Pubblicazione: (2025)
di: Peng, Tianhao, et al.
Pubblicazione: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
di: Faraz, Ali, et al.
Pubblicazione: (2025)
di: Faraz, Ali, et al.
Pubblicazione: (2025)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
di: Lian, Weitong, et al.
Pubblicazione: (2026)
di: Lian, Weitong, et al.
Pubblicazione: (2026)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
di: Wan, Hanwen, et al.
Pubblicazione: (2025)
di: Wan, Hanwen, et al.
Pubblicazione: (2025)
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
di: Singhania, Aditi, et al.
Pubblicazione: (2025)
di: Singhania, Aditi, et al.
Pubblicazione: (2025)
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
di: Zaazou, Youssef, et al.
Pubblicazione: (2026)
di: Zaazou, Youssef, et al.
Pubblicazione: (2026)
Caption This, Reason That: VLMs Caught in the Middle
di: Weng, Zihan, et al.
Pubblicazione: (2025)
di: Weng, Zihan, et al.
Pubblicazione: (2025)
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
di: Qian, Chen, et al.
Pubblicazione: (2025)
di: Qian, Chen, et al.
Pubblicazione: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
di: Singh, Aditya Kumar, et al.
Pubblicazione: (2026)
di: Singh, Aditya Kumar, et al.
Pubblicazione: (2026)
Empowering Semantic-Sensitive Underwater Image Enhancement with VLM
di: Fan, Guodong, et al.
Pubblicazione: (2026)
di: Fan, Guodong, et al.
Pubblicazione: (2026)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
VACoT: Rethinking Visual Data Augmentation with VLMs
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images
di: Sarowar, Md Selim, et al.
Pubblicazione: (2025)
di: Sarowar, Md Selim, et al.
Pubblicazione: (2025)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
di: Maniyar, Suyash, et al.
Pubblicazione: (2025)
di: Maniyar, Suyash, et al.
Pubblicazione: (2025)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
di: Zhang, Yiman, et al.
Pubblicazione: (2025)
di: Zhang, Yiman, et al.
Pubblicazione: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
di: Li, Frank, et al.
Pubblicazione: (2025) -
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2025) -
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
di: Wang, Shengao, et al.
Pubblicazione: (2025) -
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025) -
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
di: Kang, Inha, et al.
Pubblicazione: (2025)