Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zejian, Li, Yize, Meng, Chenye, Liu, Zhongni, Ling, Yang, Zhang, Shengyuan, Yang, Guang, Yang, Changyuan, Yang, Zhiyuan, Sun, Lingyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911087805333504
author Li, Zejian
Li, Yize
Meng, Chenye
Liu, Zhongni
Ling, Yang
Zhang, Shengyuan
Yang, Guang
Yang, Changyuan
Yang, Zhiyuan
Sun, Lingyun
author_facet Li, Zejian
Li, Yize
Meng, Chenye
Liu, Zhongni
Ling, Yang
Zhang, Shengyuan
Yang, Guang
Yang, Changyuan
Yang, Zhiyuan
Sun, Lingyun
contents Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO
format Preprint
id arxiv_https___arxiv_org_abs_2507_11554
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
Li, Zejian
Li, Yize
Meng, Chenye
Liu, Zhongni
Ling, Yang
Zhang, Shengyuan
Yang, Guang
Yang, Changyuan
Yang, Zhiyuan
Sun, Lingyun
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO
title Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.11554