UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jingwei, Wu, Ruoxi, Shen, Wei, Li, Meng, Liu, Yulong, She, Huimin, Yuan, Lunxi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944944631808
author Yang, Jingwei
Wu, Ruoxi
Shen, Wei
Li, Meng
Liu, Yulong
She, Huimin
Yuan, Lunxi
author_facet Yang, Jingwei
Wu, Ruoxi
Shen, Wei
Li, Meng
Liu, Yulong
She, Huimin
Yuan, Lunxi
contents Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, leading to reference-content leakage and unstable generation. We present UniCSG, a unified framework for content-constrained, style-driven generation in both text-guided and reference-guided settings. UniCSG employs staged training: (i) a latent-space semantic disentanglement stage that combines low-frequency preprocessing with conditioning corruption to encourage content-style separation, and (ii) a latent-space frequency-aware detail reconstruction stage that refines details via multi-scale frequency supervision. We further incorporate pixel-space reward learning to align latent objectives with perceptual quality after decoding. Experiments demonstrate improved content faithfulness, style alignment, and robustness in both settings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17850
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
Yang, Jingwei
Wu, Ruoxi
Shen, Wei
Li, Meng
Liu, Yulong
She, Huimin
Yuan, Lunxi
Computer Vision and Pattern Recognition
Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, leading to reference-content leakage and unstable generation. We present UniCSG, a unified framework for content-constrained, style-driven generation in both text-guided and reference-guided settings. UniCSG employs staged training: (i) a latent-space semantic disentanglement stage that combines low-frequency preprocessing with conditioning corruption to encourage content-style separation, and (ii) a latent-space frequency-aware detail reconstruction stage that refines details via multi-scale frequency supervision. We further incorporate pixel-space reward learning to align latent objectives with perceptual quality after decoding. Experiments demonstrate improved content faithfulness, style alignment, and robustness in both settings.
title UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.17850