MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Huiyi, Peng, Jiawei, Min, Dehai, Sun, Changchang, Chen, Kaijie, Yan, Yan, Yang, Xu, Cheng, Lu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
von: Yan, Bei, et al.
Veröffentlicht: (2025)
von: Yan, Bei, et al.
Veröffentlicht: (2025)
Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
von: Sui, Xiangjie, et al.
Veröffentlicht: (2025)
von: Sui, Xiangjie, et al.
Veröffentlicht: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance
von: Lin, Chenwei, et al.
Veröffentlicht: (2024)
von: Lin, Chenwei, et al.
Veröffentlicht: (2024)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Visual Grounding with Attention-Driven Constraint Balancing
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
von: Zhang, Yifei, et al.
Veröffentlicht: (2023)
von: Zhang, Yifei, et al.
Veröffentlicht: (2023)
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
von: Zhang, Haozhe, et al.
Veröffentlicht: (2026)
von: Zhang, Haozhe, et al.
Veröffentlicht: (2026)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
von: Zhuang, Cailin, et al.
Veröffentlicht: (2025)
von: Zhuang, Cailin, et al.
Veröffentlicht: (2025)
M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
von: Weng, Ju-Hsuan, et al.
Veröffentlicht: (2025)
von: Weng, Ju-Hsuan, et al.
Veröffentlicht: (2025)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
von: Sun, Changchang, et al.
Veröffentlicht: (2024)
von: Sun, Changchang, et al.
Veröffentlicht: (2024)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
Can GPT-4 Models Detect Misleading Visualizations?
von: Alexander, Jason, et al.
Veröffentlicht: (2024)
von: Alexander, Jason, et al.
Veröffentlicht: (2024)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
Visual Text Processing: A Comprehensive Review and Unified Evaluation
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
Adaptive Visual Navigation Assistant in 3D RPGs
von: Xu, Kaijie, et al.
Veröffentlicht: (2025)
von: Xu, Kaijie, et al.
Veröffentlicht: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
von: Yan, Bei, et al.
Veröffentlicht: (2025) -
Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
von: Sui, Xiangjie, et al.
Veröffentlicht: (2025) -
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025) -
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
von: Sun, Changchang, et al.
Veröffentlicht: (2025) -
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
von: Zhang, Jie, et al.
Veröffentlicht: (2024)