AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Mengzhao, Zhang, Zhihan, Cases, Ignacio, Liu, Zheyuan, Jiang, Meng, Qi, Peng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944032370688
author Jia, Mengzhao
Zhang, Zhihan
Cases, Ignacio
Liu, Zheyuan
Jiang, Meng
Qi, Peng
author_facet Jia, Mengzhao
Zhang, Zhihan
Cases, Ignacio
Liu, Zheyuan
Jiang, Meng
Qi, Peng
contents Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Jia, Mengzhao
Zhang, Zhihan
Cases, Ignacio
Liu, Zheyuan
Jiang, Meng
Qi, Peng
Computation and Language
Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations.
title AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
topic Computation and Language
url https://arxiv.org/abs/2510.14738