SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Pingchuan, Yang, Xiaopei, Li, Yusong, Gui, Ming, Krause, Felix, Schusterbauer, Johannes, Ommer, Björn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918115366928384
author Ma, Pingchuan
Yang, Xiaopei
Li, Yusong
Gui, Ming
Krause, Felix
Schusterbauer, Johannes
Ommer, Björn
author_facet Ma, Pingchuan
Yang, Xiaopei
Li, Yusong
Gui, Ming
Krause, Felix
Schusterbauer, Johannes
Ommer, Björn
contents Explicitly disentangling style and content in vision models remains challenging due to their semantic overlap and the subjectivity of human perception. Existing methods propose separation through generative or discriminative objectives, but they still face the inherent ambiguity of disentangling intertwined concepts. Instead, we ask: Can we bypass explicit disentanglement by learning to merge style and content invertibly, allowing separation to emerge naturally? We propose SCFlow, a flow-matching framework that learns bidirectional mappings between entangled and disentangled representations. Our approach is built upon three key insights: 1) Training solely to merge style and content, a well-defined task, enables invertible disentanglement without explicit supervision; 2) flow matching bridges on arbitrary distributions, avoiding the restrictive Gaussian priors of diffusion models and normalizing flows; and 3) a synthetic dataset of 510,000 samples (51 styles $\times$ 10,000 content samples) was curated to simulate disentanglement through systematic style-content pairing. Beyond controllable generation tasks, we demonstrate that SCFlow generalizes to ImageNet-1k and WikiArt in zero-shot settings and achieves competitive performance, highlighting that disentanglement naturally emerges from the invertible merging process.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
Ma, Pingchuan
Yang, Xiaopei
Li, Yusong
Gui, Ming
Krause, Felix
Schusterbauer, Johannes
Ommer, Björn
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Explicitly disentangling style and content in vision models remains challenging due to their semantic overlap and the subjectivity of human perception. Existing methods propose separation through generative or discriminative objectives, but they still face the inherent ambiguity of disentangling intertwined concepts. Instead, we ask: Can we bypass explicit disentanglement by learning to merge style and content invertibly, allowing separation to emerge naturally? We propose SCFlow, a flow-matching framework that learns bidirectional mappings between entangled and disentangled representations. Our approach is built upon three key insights: 1) Training solely to merge style and content, a well-defined task, enables invertible disentanglement without explicit supervision; 2) flow matching bridges on arbitrary distributions, avoiding the restrictive Gaussian priors of diffusion models and normalizing flows; and 3) a synthetic dataset of 510,000 samples (51 styles $\times$ 10,000 content samples) was curated to simulate disentanglement through systematic style-content pairing. Beyond controllable generation tasks, we demonstrate that SCFlow generalizes to ImageNet-1k and WikiArt in zero-shot settings and achieves competitive performance, highlighting that disentanglement naturally emerges from the invertible merging process.
title SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.03402