Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hwang, Yerin, Lee, Dongryeol, Min, Kyungmin, Kang, Taegwan, Kim, Yong-il, Jung, Kyomin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
When Wording Steers the Evaluation: Framing Bias in LLM judges
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
di: Park, Seongheon, et al.
Pubblicazione: (2026)
di: Park, Seongheon, et al.
Pubblicazione: (2026)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
di: Gong, Xuan, et al.
Pubblicazione: (2024)
di: Gong, Xuan, et al.
Pubblicazione: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
LVLM-Composer's Explicit Planning for Image Generation
di: Ramsey, Spencer, et al.
Pubblicazione: (2025)
di: Ramsey, Spencer, et al.
Pubblicazione: (2025)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
Return of EM: Entity-driven Answer Set Expansion for QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference
di: Huang, Zhaohong, et al.
Pubblicazione: (2026)
di: Huang, Zhaohong, et al.
Pubblicazione: (2026)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
di: Wang, Teng, et al.
Pubblicazione: (2024)
di: Wang, Teng, et al.
Pubblicazione: (2024)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
di: Zhu, Lanyun, et al.
Pubblicazione: (2025)
di: Zhu, Lanyun, et al.
Pubblicazione: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
di: Lee, Sua, et al.
Pubblicazione: (2026)
di: Lee, Sua, et al.
Pubblicazione: (2026)
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
di: Yan, Bei, et al.
Pubblicazione: (2026)
di: Yan, Bei, et al.
Pubblicazione: (2026)
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
di: Feng, Ze, et al.
Pubblicazione: (2025)
di: Feng, Ze, et al.
Pubblicazione: (2025)
LLMs can be easily Confused by Instructional Distractions
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
di: Garcia, Fernando Gabriela, et al.
Pubblicazione: (2025)
di: Garcia, Fernando Gabriela, et al.
Pubblicazione: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
di: Ahn, Sunghyun, et al.
Pubblicazione: (2025)
di: Ahn, Sunghyun, et al.
Pubblicazione: (2025)
AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
di: Zhong, Li'an, et al.
Pubblicazione: (2026)
di: Zhong, Li'an, et al.
Pubblicazione: (2026)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
di: Peng, Liyang, et al.
Pubblicazione: (2025)
di: Peng, Liyang, et al.
Pubblicazione: (2025)
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
di: Liu, Qingyuan, et al.
Pubblicazione: (2025)
di: Liu, Qingyuan, et al.
Pubblicazione: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
di: Stan, Gabriela Ben Melech, et al.
Pubblicazione: (2024)
di: Stan, Gabriela Ben Melech, et al.
Pubblicazione: (2024)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
di: Cheng, Ruoxi, et al.
Pubblicazione: (2026)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2026)
Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
di: Zhou, Minghao, et al.
Pubblicazione: (2025)
di: Zhou, Minghao, et al.
Pubblicazione: (2025)
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
di: Guo, Leilei, et al.
Pubblicazione: (2025)
di: Guo, Leilei, et al.
Pubblicazione: (2025)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models
di: Mazor, Nir, et al.
Pubblicazione: (2025)
di: Mazor, Nir, et al.
Pubblicazione: (2025)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
di: Dastmalchi, Hamidreza, et al.
Pubblicazione: (2026)
di: Dastmalchi, Hamidreza, et al.
Pubblicazione: (2026)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
di: Liang, Zihan, et al.
Pubblicazione: (2025)
di: Liang, Zihan, et al.
Pubblicazione: (2025)
MedM-VL: What Makes a Good Medical LVLM?
di: Shi, Yiming, et al.
Pubblicazione: (2025)
di: Shi, Yiming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025) -
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
di: Hwang, Yerin, et al.
Pubblicazione: (2025) -
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
di: Hwang, Yerin, et al.
Pubblicazione: (2026)