SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Weiyang, Niu, Yuwei, Liao, Jiaqi, Duan, Chengqi, Li, Aoxue, Gao, Shenghua, Liu, Xihui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918160582574080
author Jin, Weiyang
Niu, Yuwei
Liao, Jiaqi
Duan, Chengqi
Li, Aoxue
Gao, Shenghua
Liu, Xihui
author_facet Jin, Weiyang
Niu, Yuwei
Liao, Jiaqi
Duan, Chengqi
Li, Aoxue
Gao, Shenghua
Liu, Xihui
contents Recently, remarkable progress has been made in Unified Multimodal Models (UMMs), which integrate vision-language generation and understanding capabilities within a single framework. However, a significant gap exists where a model's strong visual understanding often fails to transfer to its visual generation. A model might correctly understand an image based on user instructions, yet be unable to generate a faithful image from text prompts. This phenomenon directly raises a compelling question: Can a model achieve self-improvement by using its understanding module to reward its generation module? To bridge this gap and achieve self-improvement, we introduce SRUM, a self-rewarding post-training framework that can be directly applied to existing UMMs of various designs. SRUM creates a feedback loop where the model's own understanding module acts as an internal ``evaluator'', providing corrective signals to improve its generation module, without requiring additional human-labeled data. To ensure this feedback is comprehensive, we designed a global-local dual reward system. To tackle the inherent structural complexity of images, this system offers multi-scale guidance: a \textbf{global reward} ensures the correctness of the overall visual semantics and layout, while a \textbf{local reward} refines fine-grained, object-level fidelity. SRUM leads to powerful capabilities and shows strong generalization, boosting performance on T2I-CompBench from 82.18 to \textbf{88.37} and on T2I-ReasonBench from 43.82 to \textbf{46.75}. Overall, our work establishes a powerful new paradigm for enabling a UMMs' understanding module to guide and enhance its own generation via self-rewarding.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12784
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
Jin, Weiyang
Niu, Yuwei
Liao, Jiaqi
Duan, Chengqi
Li, Aoxue
Gao, Shenghua
Liu, Xihui
Computer Vision and Pattern Recognition
Computation and Language
I.4.0
Recently, remarkable progress has been made in Unified Multimodal Models (UMMs), which integrate vision-language generation and understanding capabilities within a single framework. However, a significant gap exists where a model's strong visual understanding often fails to transfer to its visual generation. A model might correctly understand an image based on user instructions, yet be unable to generate a faithful image from text prompts. This phenomenon directly raises a compelling question: Can a model achieve self-improvement by using its understanding module to reward its generation module? To bridge this gap and achieve self-improvement, we introduce SRUM, a self-rewarding post-training framework that can be directly applied to existing UMMs of various designs. SRUM creates a feedback loop where the model's own understanding module acts as an internal ``evaluator'', providing corrective signals to improve its generation module, without requiring additional human-labeled data. To ensure this feedback is comprehensive, we designed a global-local dual reward system. To tackle the inherent structural complexity of images, this system offers multi-scale guidance: a \textbf{global reward} ensures the correctness of the overall visual semantics and layout, while a \textbf{local reward} refines fine-grained, object-level fidelity. SRUM leads to powerful capabilities and shows strong generalization, boosting performance on T2I-CompBench from 82.18 to \textbf{88.37} and on T2I-ReasonBench from 43.82 to \textbf{46.75}. Overall, our work establishes a powerful new paradigm for enabling a UMMs' understanding module to guide and enhance its own generation via self-rewarding.
title SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
topic Computer Vision and Pattern Recognition
Computation and Language
I.4.0
url https://arxiv.org/abs/2510.12784