Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xian, Jia Jun Cheng, Li, Muchen, Yang, Haotian, Tao, Xin, Wan, Pengfei, Sigal, Leonid, Liao, Renjie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914067271122944
author Xian, Jia Jun Cheng
Li, Muchen
Yang, Haotian
Tao, Xin
Wan, Pengfei
Sigal, Leonid
Liao, Renjie
author_facet Xian, Jia Jun Cheng
Li, Muchen
Yang, Haotian
Tao, Xin
Wan, Pengfei
Sigal, Leonid
Liao, Renjie
contents Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alignment between the text and the generated image remains a significant challenge for state-of-the-art diffusion models. To address this, existing studies employ reinforcement learning with human feedback (RLHF) to align T2I outputs with human preferences. These methods, however, either rely directly on paired image preference data or require a learned reward function, both of which depend heavily on costly, high-quality human annotations and thus face scalability limitations. In this work, we introduce Text Preference Optimization (TPO), a framework that enables "free-lunch" alignment of T2I models, achieving alignment without the need for paired image preference data. TPO works by training the model to prefer matched prompts over mismatched prompts, which are constructed by perturbing original captions using a large language model. Our framework is general and compatible with existing preference-based algorithms. We extend both DPO and KTO to our setting, resulting in TDPO and TKTO. Quantitative and qualitative evaluations across multiple benchmarks show that our methods consistently outperform their original counterparts, delivering better human preference scores and improved text-to-image alignment. Our Open-source code is available at https://github.com/DSL-Lab/T2I-Free-Lunch-Alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25771
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
Xian, Jia Jun Cheng
Li, Muchen
Yang, Haotian
Tao, Xin
Wan, Pengfei
Sigal, Leonid
Liao, Renjie
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alignment between the text and the generated image remains a significant challenge for state-of-the-art diffusion models. To address this, existing studies employ reinforcement learning with human feedback (RLHF) to align T2I outputs with human preferences. These methods, however, either rely directly on paired image preference data or require a learned reward function, both of which depend heavily on costly, high-quality human annotations and thus face scalability limitations. In this work, we introduce Text Preference Optimization (TPO), a framework that enables "free-lunch" alignment of T2I models, achieving alignment without the need for paired image preference data. TPO works by training the model to prefer matched prompts over mismatched prompts, which are constructed by perturbing original captions using a large language model. Our framework is general and compatible with existing preference-based algorithms. We extend both DPO and KTO to our setting, resulting in TDPO and TKTO. Quantitative and qualitative evaluations across multiple benchmarks show that our methods consistently outperform their original counterparts, delivering better human preference scores and improved text-to-image alignment. Our Open-source code is available at https://github.com/DSL-Lab/T2I-Free-Lunch-Alignment.
title Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.25771