MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhai, Kevin, Singh, Utsav, Thatipelli, Anirudh, Chakraborty, Souradip, Sahu, Anit Kumar, Huang, Furong, Bedi, Amrit Singh, Shah, Mubarak
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908572958326784
author Zhai, Kevin
Singh, Utsav
Thatipelli, Anirudh
Chakraborty, Souradip
Sahu, Anit Kumar
Huang, Furong
Bedi, Amrit Singh
Shah, Mubarak
author_facet Zhai, Kevin
Singh, Utsav
Thatipelli, Anirudh
Chakraborty, Souradip
Sahu, Anit Kumar
Huang, Furong
Bedi, Amrit Singh
Shah, Mubarak
contents Diffusion models excel at generating images conditioned on text prompts, but the resulting images often do not satisfy user-specific criteria measured by scalar rewards such as Aesthetic Scores. This alignment typically requires fine-tuning, which is computationally demanding. Recently, inference-time alignment via noise optimization has emerged as an efficient alternative, modifying initial input noise to steer the diffusion denoising process towards generating high-reward images. However, this approach suffers from reward hacking, where the model produces images that score highly, yet deviate significantly from the original prompt. We show that noise-space regularization is insufficient and that preventing reward hacking requires an explicit image-space constraint. To this end, we propose MIRA (MItigating Reward hAcking), a training-free, inference-time alignment method. MIRA introduces an image-space, score-based KL surrogate that regularizes the sampling trajectory with a frozen backbone, constraining the output distribution so reward can increase without off-distribution drift (reward hacking). We derive a tractable approximation to KL using diffusion scores. Across SDv1.5 and SDXL, multiple rewards (Aesthetic, HPSv2, PickScore), and public datasets (e.g., Animal-Animal, HPDv2), MIRA achieves >60\% win rate vs. strong baselines while preserving prompt adherence; mechanism plots show reward gains with near-zero drift, whereas DNO drifts as compute increases. We further introduce MIRA-DPO, mapping preference optimization to inference time with a frozen backbone, extending MIRA to non-differentiable rewards without fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
Zhai, Kevin
Singh, Utsav
Thatipelli, Anirudh
Chakraborty, Souradip
Sahu, Anit Kumar
Huang, Furong
Bedi, Amrit Singh
Shah, Mubarak
Machine Learning
Diffusion models excel at generating images conditioned on text prompts, but the resulting images often do not satisfy user-specific criteria measured by scalar rewards such as Aesthetic Scores. This alignment typically requires fine-tuning, which is computationally demanding. Recently, inference-time alignment via noise optimization has emerged as an efficient alternative, modifying initial input noise to steer the diffusion denoising process towards generating high-reward images. However, this approach suffers from reward hacking, where the model produces images that score highly, yet deviate significantly from the original prompt. We show that noise-space regularization is insufficient and that preventing reward hacking requires an explicit image-space constraint. To this end, we propose MIRA (MItigating Reward hAcking), a training-free, inference-time alignment method. MIRA introduces an image-space, score-based KL surrogate that regularizes the sampling trajectory with a frozen backbone, constraining the output distribution so reward can increase without off-distribution drift (reward hacking). We derive a tractable approximation to KL using diffusion scores. Across SDv1.5 and SDXL, multiple rewards (Aesthetic, HPSv2, PickScore), and public datasets (e.g., Animal-Animal, HPDv2), MIRA achieves >60\% win rate vs. strong baselines while preserving prompt adherence; mechanism plots show reward gains with near-zero drift, whereas DNO drifts as compute increases. We further introduce MIRA-DPO, mapping preference optimization to inference time with a frozen backbone, extending MIRA to non-differentiable rewards without fine-tuning.
title MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2510.01549