MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911697830150144 |
|---|---|
| author | Dufour, Nicolas Degeorge, Lucas Ghosh, Arijit Kalogeiton, Vicky Picard, David |
| author_facet | Dufour, Nicolas Degeorge, Lucas Ghosh, Arijit Kalogeiton, Vicky Picard, David |
| contents | The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25897 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency Dufour, Nicolas Degeorge, Lucas Ghosh, Arijit Kalogeiton, Vicky Picard, David Computer Vision and Pattern Recognition Machine Learning The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2). |
| title | MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2510.25897 |