USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shaojin, Huang, Mengqi, Cheng, Yufeng, Wu, Wenxu, Tian, Jiahe, Luo, Yiming, Ding, Fei, He, Qian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914006369828864
author Wu, Shaojin
Huang, Mengqi
Cheng, Yufeng
Wu, Wenxu
Tian, Jiahe
Luo, Yiming
Ding, Fei
He, Qian
author_facet Wu, Shaojin
Huang, Mengqi
Cheng, Yufeng
Wu, Wenxu
Tian, Jiahe
Luo, Yiming
Ding, Fei
He, Qian
contents Existing literature typically treats style-driven and subject-driven generation as two disjoint tasks: the former prioritizes stylistic similarity, whereas the latter insists on subject consistency, resulting in an apparent antagonism. We argue that both objectives can be unified under a single framework because they ultimately concern the disentanglement and re-composition of content and style, a long-standing theme in style-driven research. To this end, we present USO, a Unified Style-Subject Optimized customization model. First, we construct a large-scale triplet dataset consisting of content images, style images, and their corresponding stylized content images. Second, we introduce a disentangled learning scheme that simultaneously aligns style features and disentangles content from style through two complementary objectives, style-alignment training and content-style disentanglement training. Third, we incorporate a style reward-learning paradigm denoted as SRL to further enhance the model's performance. Finally, we release USO-Bench, the first benchmark that jointly evaluates style similarity and subject fidelity across multiple metrics. Extensive experiments demonstrate that USO achieves state-of-the-art performance among open-source models along both dimensions of subject consistency and style similarity. Code and model: https://github.com/bytedance/USO
format Preprint
id arxiv_https___arxiv_org_abs_2508_18966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
Wu, Shaojin
Huang, Mengqi
Cheng, Yufeng
Wu, Wenxu
Tian, Jiahe
Luo, Yiming
Ding, Fei
He, Qian
Computer Vision and Pattern Recognition
Machine Learning
Existing literature typically treats style-driven and subject-driven generation as two disjoint tasks: the former prioritizes stylistic similarity, whereas the latter insists on subject consistency, resulting in an apparent antagonism. We argue that both objectives can be unified under a single framework because they ultimately concern the disentanglement and re-composition of content and style, a long-standing theme in style-driven research. To this end, we present USO, a Unified Style-Subject Optimized customization model. First, we construct a large-scale triplet dataset consisting of content images, style images, and their corresponding stylized content images. Second, we introduce a disentangled learning scheme that simultaneously aligns style features and disentangles content from style through two complementary objectives, style-alignment training and content-style disentanglement training. Third, we incorporate a style reward-learning paradigm denoted as SRL to further enhance the model's performance. Finally, we release USO-Bench, the first benchmark that jointly evaluates style similarity and subject fidelity across multiple metrics. Extensive experiments demonstrate that USO achieves state-of-the-art performance among open-source models along both dimensions of subject consistency and style similarity. Code and model: https://github.com/bytedance/USO
title USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.18966