Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cohen, Nadav Z., Nir, Oron, Shamir, Ariel
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908476200976384
author Cohen, Nadav Z.
Nir, Oron
Shamir, Ariel
author_facet Cohen, Nadav Z.
Nir, Oron
Shamir, Ariel
contents Balancing content fidelity and artistic style is a pivotal challenge in image generation. While traditional style transfer methods and modern Denoising Diffusion Probabilistic Models (DDPMs) strive to achieve this balance, they often struggle to do so without sacrificing either style, content, or sometimes both. This work addresses this challenge by analyzing the ability of DDPMs to maintain content and style equilibrium. We introduce a novel method to identify sensitivities within the DDPM attention layers, identifying specific layers that correspond to different stylistic aspects. By directing conditional inputs only to these sensitive layers, our approach enables fine-grained control over style and content, significantly reducing issues arising from over-constrained inputs. Our findings demonstrate that this method enhances recent stylization techniques by better aligning style and content, ultimately improving the quality of generated visual content.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19853
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation
Cohen, Nadav Z.
Nir, Oron
Shamir, Ariel
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Balancing content fidelity and artistic style is a pivotal challenge in image generation. While traditional style transfer methods and modern Denoising Diffusion Probabilistic Models (DDPMs) strive to achieve this balance, they often struggle to do so without sacrificing either style, content, or sometimes both. This work addresses this challenge by analyzing the ability of DDPMs to maintain content and style equilibrium. We introduce a novel method to identify sensitivities within the DDPM attention layers, identifying specific layers that correspond to different stylistic aspects. By directing conditional inputs only to these sensitive layers, our approach enables fine-grained control over style and content, significantly reducing issues arising from over-constrained inputs. Our findings demonstrate that this method enhances recent stylization techniques by better aligning style and content, ultimately improving the quality of generated visual content.
title Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2412.19853