Style-Friendly SNR Sampler for Style-Driven Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Jooyoung, Shin, Chaehun, Oh, Yeongtak, Kim, Heeseung, Lee, Jungbeom, Yoon, Sungroh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909544426242048
author Choi, Jooyoung
Shin, Chaehun
Oh, Yeongtak
Kim, Heeseung
Lee, Jungbeom
Yoon, Sungroh
author_facet Choi, Jooyoung
Shin, Chaehun
Oh, Yeongtak
Kim, Heeseung
Lee, Jungbeom
Yoon, Sungroh
contents Recent text-to-image diffusion models generate high-quality images but struggle to learn new, personalized styles, which limits the creation of unique style templates. In style-driven generation, users typically supply reference images exemplifying the desired style, together with text prompts that specify desired stylistic attributes. Previous approaches popularly rely on fine-tuning, yet it often blindly utilizes objectives and noise level distributions from pre-training without adaptation. We discover that stylistic features predominantly emerge at higher noise levels, leading current fine-tuning methods to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which aggressively shifts the signal-to-noise ratio (SNR) distribution toward higher noise levels during fine-tuning to focus on noise levels where stylistic features emerge. This enhances models' ability to capture novel styles indicated by reference images and text prompts. We demonstrate improved generation of novel styles that cannot be adequately described solely with a text prompt, enabling the creation of new style templates for personalized content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14793
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Style-Friendly SNR Sampler for Style-Driven Generation
Choi, Jooyoung
Shin, Chaehun
Oh, Yeongtak
Kim, Heeseung
Lee, Jungbeom
Yoon, Sungroh
Computer Vision and Pattern Recognition
Recent text-to-image diffusion models generate high-quality images but struggle to learn new, personalized styles, which limits the creation of unique style templates. In style-driven generation, users typically supply reference images exemplifying the desired style, together with text prompts that specify desired stylistic attributes. Previous approaches popularly rely on fine-tuning, yet it often blindly utilizes objectives and noise level distributions from pre-training without adaptation. We discover that stylistic features predominantly emerge at higher noise levels, leading current fine-tuning methods to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which aggressively shifts the signal-to-noise ratio (SNR) distribution toward higher noise levels during fine-tuning to focus on noise levels where stylistic features emerge. This enhances models' ability to capture novel styles indicated by reference images and text prompts. We demonstrate improved generation of novel styles that cannot be adequately described solely with a text prompt, enabling the creation of new style templates for personalized content creation.
title Style-Friendly SNR Sampler for Style-Driven Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.14793