AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915944032370688 |
|---|---|
| author | Jia, Mengzhao Zhang, Zhihan Cases, Ignacio Liu, Zheyuan Jiang, Meng Qi, Peng |
| author_facet | Jia, Mengzhao Zhang, Zhihan Cases, Ignacio Liu, Zheyuan Jiang, Meng Qi, Peng |
| contents | Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_14738 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning Jia, Mengzhao Zhang, Zhihan Cases, Ignacio Liu, Zheyuan Jiang, Meng Qi, Peng Computation and Language Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations. |
| title | AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.14738 |