RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Xuelu, Li, Yunsheng, Wan, Ziyu, Gao, Zixuan, Yuan, Junsong, Chen, Dongdong, Qiao, Chunming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909924928258048
author Feng, Xuelu
Li, Yunsheng
Wan, Ziyu
Gao, Zixuan
Yuan, Junsong
Chen, Dongdong
Qiao, Chunming
author_facet Feng, Xuelu
Li, Yunsheng
Wan, Ziyu
Gao, Zixuan
Yuan, Junsong
Chen, Dongdong
Qiao, Chunming
contents Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in designing effective and interpretable rewards. Existing methods often rely on either composite metrics (e.g., CLIP, OCR, and realism scores) with fixed weights or a single scalar reward distilled from human preference models, which can limit interpretability and flexibility. We propose RubricRL, a simple and general framework for rubric-based reward design that offers greater interpretability, composability, and user control. Instead of using a black-box scalar signal, RubricRL dynamically constructs a structured rubric for each prompt--a decomposable checklist of fine-grained visual criteria such as object correctness, attribute accuracy, OCR fidelity, and realism--tailored to the input text. Each criterion is independently evaluated by a multimodal judge (e.g., o4-mini), and a prompt-adaptive weighting mechanism emphasizes the most relevant dimensions. This design not only produces interpretable and modular supervision signals for policy optimization (e.g., GRPO or PPO), but also enables users to directly adjust which aspects to reward or penalize. Experiments with an autoregressive text-to-image model demonstrate that RubricRL improves prompt faithfulness, visual detail, and generalizability, while offering a flexible and extensible foundation for interpretable RL alignment across text-to-image architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20651
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
Feng, Xuelu
Li, Yunsheng
Wan, Ziyu
Gao, Zixuan
Yuan, Junsong
Chen, Dongdong
Qiao, Chunming
Computer Vision and Pattern Recognition
Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in designing effective and interpretable rewards. Existing methods often rely on either composite metrics (e.g., CLIP, OCR, and realism scores) with fixed weights or a single scalar reward distilled from human preference models, which can limit interpretability and flexibility. We propose RubricRL, a simple and general framework for rubric-based reward design that offers greater interpretability, composability, and user control. Instead of using a black-box scalar signal, RubricRL dynamically constructs a structured rubric for each prompt--a decomposable checklist of fine-grained visual criteria such as object correctness, attribute accuracy, OCR fidelity, and realism--tailored to the input text. Each criterion is independently evaluated by a multimodal judge (e.g., o4-mini), and a prompt-adaptive weighting mechanism emphasizes the most relevant dimensions. This design not only produces interpretable and modular supervision signals for policy optimization (e.g., GRPO or PPO), but also enables users to directly adjust which aspects to reward or penalize. Experiments with an autoregressive text-to-image model demonstrate that RubricRL improves prompt faithfulness, visual detail, and generalizability, while offering a flexible and extensible foundation for interpretable RL alignment across text-to-image architectures.
title RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20651