Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Chenye, Li, Zejian, Liu, Zhongni, Li, Yize, Xie, Changle, Jia, Kaixin, Yang, Ling, Deng, Huanghuang, Ding, Shiying, Zhang, Shengyuan, Li, Jiayi, Sun, Lingyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911360043974656
author Meng, Chenye
Li, Zejian
Liu, Zhongni
Li, Yize
Xie, Changle
Jia, Kaixin
Yang, Ling
Deng, Huanghuang
Ding, Shiying
Zhang, Shengyuan
Li, Jiayi
Sun, Lingyun
author_facet Meng, Chenye
Li, Zejian
Liu, Zhongni
Li, Yize
Xie, Changle
Jia, Kaixin
Yang, Ling
Deng, Huanghuang
Ding, Shiying
Zhang, Shengyuan
Li, Jiayi
Sun, Lingyun
contents Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is hierarchical and fine-grained. To address this, we first construct a hierarchical, fine-grained evaluation criteria with domain experts, which decomposes image quality into multiple positive and negative attributes organized in a tree structure. Building on this, we propose a two-stage alignment framework. First, we inject domain knowledge to an auxiliary diffusion model via Supervised Fine-Tuning. Second, we introduce Complex Preference Optimization (CPO) that extends DPO to align the target diffusion to our non-binary, hierarchical criteria. Specifically, we reformulate the alignment problem to simultaneously maximize the probability of positive attributes while minimizing the probability of negative attributes with the auxiliary diffusion. We instantiate our approach in the domain of painting generation and conduct CPO training with an annotated dataset of painting with fine-grained attributes based on our criteria. Extensive experiments demonstrate that CPO significantly enhances generation quality and alignment with expertise, opening new avenues for fine-grained criteria alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04300
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes
Meng, Chenye
Li, Zejian
Liu, Zhongni
Li, Yize
Xie, Changle
Jia, Kaixin
Yang, Ling
Deng, Huanghuang
Ding, Shiying
Zhang, Shengyuan
Li, Jiayi
Sun, Lingyun
Computer Vision and Pattern Recognition
Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is hierarchical and fine-grained. To address this, we first construct a hierarchical, fine-grained evaluation criteria with domain experts, which decomposes image quality into multiple positive and negative attributes organized in a tree structure. Building on this, we propose a two-stage alignment framework. First, we inject domain knowledge to an auxiliary diffusion model via Supervised Fine-Tuning. Second, we introduce Complex Preference Optimization (CPO) that extends DPO to align the target diffusion to our non-binary, hierarchical criteria. Specifically, we reformulate the alignment problem to simultaneously maximize the probability of positive attributes while minimizing the probability of negative attributes with the auxiliary diffusion. We instantiate our approach in the domain of painting generation and conduct CPO training with an annotated dataset of painting with fine-grained attributes based on our criteria. Extensive experiments demonstrate that CPO significantly enhances generation quality and alignment with expertise, opening new avenues for fine-grained criteria alignment.
title Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.04300