MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dufour, Nicolas, Degeorge, Lucas, Ghosh, Arijit, Kalogeiton, Vicky, Picard, David
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911697830150144
author Dufour, Nicolas
Degeorge, Lucas
Ghosh, Arijit
Kalogeiton, Vicky
Picard, David
author_facet Dufour, Nicolas
Degeorge, Lucas
Ghosh, Arijit
Kalogeiton, Vicky
Picard, David
contents The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).
format Preprint
id arxiv_https___arxiv_org_abs_2510_25897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Dufour, Nicolas
Degeorge, Lucas
Ghosh, Arijit
Kalogeiton, Vicky
Picard, David
Computer Vision and Pattern Recognition
Machine Learning
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).
title MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.25897