Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Narayanan, Aravind, Khazaie, Vahid Reza, Raza, Shaina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
by: Raval, Ananya, et al.
Published: (2025)
by: Raval, Ananya, et al.
Published: (2025)
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Can Generative Models Improve Self-Supervised Representation Learning?
by: Ayromlou, Sana, et al.
Published: (2024)
by: Ayromlou, Sana, et al.
Published: (2024)
Enhancing Anomaly Detection Generalization through Knowledge Exposure: The Dual Effects of Augmentation
by: Anvari, Mohammad Akhavan, et al.
Published: (2024)
by: Anvari, Mohammad Akhavan, et al.
Published: (2024)
Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning
by: Roy, Shuvendu, et al.
Published: (2024)
by: Roy, Shuvendu, et al.
Published: (2024)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
by: Tsaprazlis, Efthymios, et al.
Published: (2026)
by: Tsaprazlis, Efthymios, et al.
Published: (2026)
Shape Bias and Robustness Evaluation via Cue Decomposition for Image Classification and Segmentation
by: Heinert, Edgar, et al.
Published: (2025)
by: Heinert, Edgar, et al.
Published: (2025)
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
by: Radwan, Ahmed Y., et al.
Published: (2026)
by: Radwan, Ahmed Y., et al.
Published: (2026)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
by: Malakouti, Sina, et al.
Published: (2024)
by: Malakouti, Sina, et al.
Published: (2024)
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
by: Toibazar, Daulet, et al.
Published: (2025)
by: Toibazar, Daulet, et al.
Published: (2025)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
by: Munje, Michael J., et al.
Published: (2025)
by: Munje, Michael J., et al.
Published: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
by: Lee, Sua, et al.
Published: (2026)
by: Lee, Sua, et al.
Published: (2026)
When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
by: Cao, Jinjin, et al.
Published: (2025)
by: Cao, Jinjin, et al.
Published: (2025)
Addressing Bias in VLMs for Glaucoma Detection Without Protected Attribute Supervision
by: Akash, Ahsan Habib, et al.
Published: (2025)
by: Akash, Ahsan Habib, et al.
Published: (2025)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
MLLM-as-a-Judge Exhibits Model Preference Bias
by: Koyama, Shuitsu, et al.
Published: (2026)
by: Koyama, Shuitsu, et al.
Published: (2026)
DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs
by: Pan, Jiazhen, et al.
Published: (2026)
by: Pan, Jiazhen, et al.
Published: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
by: Vo, Hao, et al.
Published: (2026)
by: Vo, Hao, et al.
Published: (2026)
BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
Should VLMs be Pre-trained with Image Data?
by: Keh, Sedrick, et al.
Published: (2025)
by: Keh, Sedrick, et al.
Published: (2025)
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs
by: Fang, Zhengru, et al.
Published: (2026)
by: Fang, Zhengru, et al.
Published: (2026)
Three Forensic Cues for JPEG AI Images
by: Bergmann, Sandra, et al.
Published: (2025)
by: Bergmann, Sandra, et al.
Published: (2025)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
by: Luo, Sha, et al.
Published: (2026)
by: Luo, Sha, et al.
Published: (2026)
Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
by: Liang, Xiaoye, et al.
Published: (2025)
by: Liang, Xiaoye, et al.
Published: (2025)
Benchmarking Joint Face Spoofing and Forgery Detection with Visual and Physiological Cues
by: Yu, Zitong, et al.
Published: (2022)
by: Yu, Zitong, et al.
Published: (2022)
Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
by: Ma, Wen, et al.
Published: (2026)
by: Ma, Wen, et al.
Published: (2026)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
by: Jadhav, Avadhoot, et al.
Published: (2025)
by: Jadhav, Avadhoot, et al.
Published: (2025)
Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
by: Khan, Sumra, et al.
Published: (2026)
by: Khan, Sumra, et al.
Published: (2026)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
by: Shin, Philip Wootaek, et al.
Published: (2024)
by: Shin, Philip Wootaek, et al.
Published: (2024)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
by: Aryashad, Ardalan, et al.
Published: (2025)
by: Aryashad, Ardalan, et al.
Published: (2025)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
by: Bandraupalli, Srihari, et al.
Published: (2025)
by: Bandraupalli, Srihari, et al.
Published: (2025)
Similar Items
-
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
by: Raval, Ananya, et al.
Published: (2025) -
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
by: Raza, Shaina, et al.
Published: (2025) -
Can Generative Models Improve Self-Supervised Representation Learning?
by: Ayromlou, Sana, et al.
Published: (2024) -
Enhancing Anomaly Detection Generalization through Knowledge Exposure: The Dual Effects of Augmentation
by: Anvari, Mohammad Akhavan, et al.
Published: (2024) -
Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning
by: Roy, Shuvendu, et al.
Published: (2024)