Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
Fuente:
arXiv
Saved in:
| Main Authors: | Slyman, Eric, Tanjim, Mehrab, Kafle, Kushal, Lee, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
by: Lyu, Hanjia, et al.
Published: (2024)
by: Lyu, Hanjia, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
by: Gao, Timin, et al.
Published: (2024)
by: Gao, Timin, et al.
Published: (2024)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)
by: Wu, Qiucheng, et al.
Published: (2026)
Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks
by: Yang, Zijiao, et al.
Published: (2024)
by: Yang, Zijiao, et al.
Published: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
Revisiting Multi-Modal LLM Evaluation
by: Lu, Jian, et al.
Published: (2024)
by: Lu, Jian, et al.
Published: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Judge Anything: MLLM as a Judge Across Any Modality
by: Pu, Shu, et al.
Published: (2025)
by: Pu, Shu, et al.
Published: (2025)
Dual-branch Prompting for Multimodal Machine Translation
by: Wang, Jie, et al.
Published: (2025)
by: Wang, Jie, et al.
Published: (2025)
Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
by: Kriz, Anita, et al.
Published: (2025)
by: Kriz, Anita, et al.
Published: (2025)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
by: Yang, Yuwei, et al.
Published: (2025)
by: Yang, Yuwei, et al.
Published: (2025)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
by: Liu, Zhaochen, et al.
Published: (2026)
by: Liu, Zhaochen, et al.
Published: (2026)
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
by: Lee, Sua, et al.
Published: (2026)
by: Lee, Sua, et al.
Published: (2026)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
by: Mi, Hongze, et al.
Published: (2024)
by: Mi, Hongze, et al.
Published: (2024)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
by: Shekhar, Shivanshu, et al.
Published: (2026)
by: Shekhar, Shivanshu, et al.
Published: (2026)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
by: Akdemir, Kiymet, et al.
Published: (2025)
by: Akdemir, Kiymet, et al.
Published: (2025)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
by: Jia, Sihang, et al.
Published: (2026)
by: Jia, Sihang, et al.
Published: (2026)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
by: Morbiato, Filippo, et al.
Published: (2025)
by: Morbiato, Filippo, et al.
Published: (2025)
LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation
by: Yang, Cunyuan, et al.
Published: (2026)
by: Yang, Cunyuan, et al.
Published: (2026)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
by: Diesendruck, Maurice, et al.
Published: (2024)
by: Diesendruck, Maurice, et al.
Published: (2024)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
by: Liu, Xinyang, et al.
Published: (2023)
by: Liu, Xinyang, et al.
Published: (2023)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)
by: Xu, Runsen, et al.
Published: (2025)
Similar Items
-
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024) -
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
by: Lyu, Hanjia, et al.
Published: (2024) -
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020) -
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022) -
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)