The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zichao, Wen, Xueru, Lou, Jie, Ji, Yuqiu, Lu, Yaojie, Han, Xianpei, Zhang, Debing, Sun, Le
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912385778843648
author Li, Zichao
Wen, Xueru
Lou, Jie
Ji, Yuqiu
Lu, Yaojie
Han, Xianpei
Zhang, Debing
Sun, Le
author_facet Li, Zichao
Wen, Xueru
Lou, Jie
Ji, Yuqiu
Lu, Yaojie
Han, Xianpei
Zhang, Debing
Sun, Le
contents Multimodal Reward Models (MM-RMs) are crucial for aligning Large Language Models (LLMs) with human preferences, particularly as LLMs increasingly interact with multimodal data. However, we find that MM-RMs trained on existing datasets often struggle to generalize to out-of-distribution data due to their reliance on unimodal spurious correlations, primarily text-only shortcuts within the training distribution, which prevents them from leveraging true multimodal reward functions. To address this, we introduce a Shortcut-aware MM-RM learning algorithm that mitigates this issue by dynamically reweighting training samples, shifting the distribution toward better multimodal understanding, and reducing dependence on unimodal spurious correlations. Our experiments demonstrate significant improvements in generalization, downstream task performance, and scalability, establishing a more robust framework for multimodal reward modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
Li, Zichao
Wen, Xueru
Lou, Jie
Ji, Yuqiu
Lu, Yaojie
Han, Xianpei
Zhang, Debing
Sun, Le
Computation and Language
Artificial Intelligence
Multimodal Reward Models (MM-RMs) are crucial for aligning Large Language Models (LLMs) with human preferences, particularly as LLMs increasingly interact with multimodal data. However, we find that MM-RMs trained on existing datasets often struggle to generalize to out-of-distribution data due to their reliance on unimodal spurious correlations, primarily text-only shortcuts within the training distribution, which prevents them from leveraging true multimodal reward functions. To address this, we introduce a Shortcut-aware MM-RM learning algorithm that mitigates this issue by dynamically reweighting training samples, shifting the distribution toward better multimodal understanding, and reducing dependence on unimodal spurious correlations. Our experiments demonstrate significant improvements in generalization, downstream task performance, and scalability, establishing a more robust framework for multimodal reward modeling.
title The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.03122