DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Xin, Gong, Dong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913755851390976
author Xie, Xin
Gong, Dong
author_facet Xie, Xin
Gong, Dong
contents Text-to-image diffusion model alignment is critical for improving the alignment between the generated images and human preferences. While training-based methods are constrained by high computational costs and dataset requirements, training-free alignment methods remain underexplored and are often limited by inaccurate guidance. We propose a plug-and-play training-free alignment method, DyMO, for aligning the generated images and human preferences during inference. Apart from text-aware human preference scores, we introduce a semantic alignment objective for enhancing the semantic alignment in the early stages of diffusion, relying on the fact that the attention maps are effective reflections of the semantics in noisy images. We propose dynamic scheduling of multiple objectives and intermediate recurrent steps to reflect the requirements at different steps. Experiments with diverse pre-trained diffusion models and metrics demonstrate the effectiveness and robustness of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00759
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling
Xie, Xin
Gong, Dong
Computer Vision and Pattern Recognition
Text-to-image diffusion model alignment is critical for improving the alignment between the generated images and human preferences. While training-based methods are constrained by high computational costs and dataset requirements, training-free alignment methods remain underexplored and are often limited by inaccurate guidance. We propose a plug-and-play training-free alignment method, DyMO, for aligning the generated images and human preferences during inference. Apart from text-aware human preference scores, we introduce a semantic alignment objective for enhancing the semantic alignment in the early stages of diffusion, relying on the fact that the attention maps are effective reflections of the semantics in noisy images. We propose dynamic scheduling of multiple objectives and intermediate recurrent steps to reflect the requirements at different steps. Experiments with diverse pre-trained diffusion models and metrics demonstrate the effectiveness and robustness of the proposed method.
title DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00759