Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hwang, Yerin, Lee, Dongryeol, Min, Kyungmin, Kang, Taegwan, Kim, Yong-il, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
par: Min, Kyungmin, et autres
Publié: (2024)
par: Min, Kyungmin, et autres
Publié: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
par: Moon, Jiwon, et autres
Publié: (2025)
par: Moon, Jiwon, et autres
Publié: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
par: Hwang, Yerin, et autres
Publié: (2025)
par: Hwang, Yerin, et autres
Publié: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
par: Lee, Dongryeol, et autres
Publié: (2026)
par: Lee, Dongryeol, et autres
Publié: (2026)
When Wording Steers the Evaluation: Framing Bias in LLM judges
par: Hwang, Yerin, et autres
Publié: (2026)
par: Hwang, Yerin, et autres
Publié: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
par: Lee, Kang-il, et autres
Publié: (2024)
par: Lee, Kang-il, et autres
Publié: (2024)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
par: Lee, Dongryeol, et autres
Publié: (2024)
par: Lee, Dongryeol, et autres
Publié: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
par: Min, Kyungmin, et autres
Publié: (2026)
par: Min, Kyungmin, et autres
Publié: (2026)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
par: Park, Seongheon, et autres
Publié: (2026)
par: Park, Seongheon, et autres
Publié: (2026)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
par: Cho, Beomsik, et autres
Publié: (2025)
par: Cho, Beomsik, et autres
Publié: (2025)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
par: Gong, Xuan, et autres
Publié: (2024)
par: Gong, Xuan, et autres
Publié: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
par: Kundu, Souvik, et autres
Publié: (2025)
par: Kundu, Souvik, et autres
Publié: (2025)
LVLM-Composer's Explicit Planning for Image Generation
par: Ramsey, Spencer, et autres
Publié: (2025)
par: Ramsey, Spencer, et autres
Publié: (2025)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
par: Kwon, JuneHyoung, et autres
Publié: (2026)
par: Kwon, JuneHyoung, et autres
Publié: (2026)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
par: Deng, Jingyuan, et autres
Publié: (2025)
par: Deng, Jingyuan, et autres
Publié: (2025)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
par: Koo, Jahyun, et autres
Publié: (2024)
par: Koo, Jahyun, et autres
Publié: (2024)
Return of EM: Entity-driven Answer Set Expansion for QA Evaluation
par: Lee, Dongryeol, et autres
Publié: (2024)
par: Lee, Dongryeol, et autres
Publié: (2024)
ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference
par: Huang, Zhaohong, et autres
Publié: (2026)
par: Huang, Zhaohong, et autres
Publié: (2026)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
par: Wang, Teng, et autres
Publié: (2024)
par: Wang, Teng, et autres
Publié: (2024)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
par: Zhu, Lanyun, et autres
Publié: (2025)
par: Zhu, Lanyun, et autres
Publié: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
par: Lee, Sua, et autres
Publié: (2026)
par: Lee, Sua, et autres
Publié: (2026)
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
par: Yan, Bei, et autres
Publié: (2026)
par: Yan, Bei, et autres
Publié: (2026)
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
par: Feng, Ze, et autres
Publié: (2025)
par: Feng, Ze, et autres
Publié: (2025)
LLMs can be easily Confused by Instructional Distractions
par: Hwang, Yerin, et autres
Publié: (2025)
par: Hwang, Yerin, et autres
Publié: (2025)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
par: Zhao, Xiaohan, et autres
Publié: (2026)
par: Zhao, Xiaohan, et autres
Publié: (2026)
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
par: Garcia, Fernando Gabriela, et autres
Publié: (2025)
par: Garcia, Fernando Gabriela, et autres
Publié: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
par: Ahn, Sunghyun, et autres
Publié: (2025)
par: Ahn, Sunghyun, et autres
Publié: (2025)
AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
par: Zhong, Li'an, et autres
Publié: (2026)
par: Zhong, Li'an, et autres
Publié: (2026)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
par: Peng, Liyang, et autres
Publié: (2025)
par: Peng, Liyang, et autres
Publié: (2025)
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
par: Liu, Qingyuan, et autres
Publié: (2025)
par: Liu, Qingyuan, et autres
Publié: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
par: Cheng, Ruoxi, et autres
Publié: (2026)
par: Cheng, Ruoxi, et autres
Publié: (2026)
Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
par: Zhou, Minghao, et autres
Publié: (2025)
par: Zhou, Minghao, et autres
Publié: (2025)
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
par: Guo, Leilei, et autres
Publié: (2025)
par: Guo, Leilei, et autres
Publié: (2025)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
par: Yang, Nakyeong, et autres
Publié: (2023)
par: Yang, Nakyeong, et autres
Publié: (2023)
LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models
par: Mazor, Nir, et autres
Publié: (2025)
par: Mazor, Nir, et autres
Publié: (2025)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
par: Fan, Zhiyuan, et autres
Publié: (2025)
par: Fan, Zhiyuan, et autres
Publié: (2025)
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
par: Dastmalchi, Hamidreza, et autres
Publié: (2026)
par: Dastmalchi, Hamidreza, et autres
Publié: (2026)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
par: Liang, Zihan, et autres
Publié: (2025)
par: Liang, Zihan, et autres
Publié: (2025)
MedM-VL: What Makes a Good Medical LVLM?
par: Shi, Yiming, et autres
Publié: (2025)
par: Shi, Yiming, et autres
Publié: (2025)
Documents similaires
-
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
par: Min, Kyungmin, et autres
Publié: (2024) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
par: Moon, Jiwon, et autres
Publié: (2025) -
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
par: Hwang, Yerin, et autres
Publié: (2025) -
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
par: Lee, Dongryeol, et autres
Publié: (2026) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
par: Hwang, Yerin, et autres
Publié: (2026)