Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reddy, Sanjana, Malhi, Ishaan, Ma, Sally, Dutta, Praneet
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912912630611968
author Reddy, Sanjana
Malhi, Ishaan
Ma, Sally
Dutta, Praneet
author_facet Reddy, Sanjana
Malhi, Ishaan
Ma, Sally
Dutta, Praneet
contents Existing methods for preference tuning of text-to-image (T2I) diffusion models often rely on computationally expensive generation steps to create positive and negative pairs of images. These approaches frequently yield training pairs that either lack meaningful differences, are expensive to sample and filter, or exhibit significant variance in irrelevant pixel regions, thereby degrading training efficiency. To address these limitations, we introduce "Di3PO", a novel method for constructing positive and negative pairs that isolates specific regions targeted for improvement during preference tuning, while keeping the surrounding context in the image stable. We demonstrate the efficacy of our approach by applying it to the challenging task of text rendering in diffusion models, showcasing improvements over baseline methods of SFT and DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06355
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation
Reddy, Sanjana
Malhi, Ishaan
Ma, Sally
Dutta, Praneet
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing methods for preference tuning of text-to-image (T2I) diffusion models often rely on computationally expensive generation steps to create positive and negative pairs of images. These approaches frequently yield training pairs that either lack meaningful differences, are expensive to sample and filter, or exhibit significant variance in irrelevant pixel regions, thereby degrading training efficiency. To address these limitations, we introduce "Di3PO", a novel method for constructing positive and negative pairs that isolates specific regions targeted for improvement during preference tuning, while keeping the surrounding context in the image stable. We demonstrate the efficacy of our approach by applying it to the challenging task of text rendering in diffusion models, showcasing improvements over baseline methods of SFT and DPO.
title Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.06355