TCFG: Tangential Damping Classifier-free Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Mingi, Kim, Shin seong, Hsiao, Jaeseok Jeong. Yi Ting, Uh, Youngjung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910890027122688
author Kwon, Mingi
Kim, Shin seong
Hsiao, Jaeseok Jeong. Yi Ting
Uh, Youngjung
author_facet Kwon, Mingi
Kim, Shin seong
Hsiao, Jaeseok Jeong. Yi Ting
Uh, Youngjung
contents Diffusion models have achieved remarkable success in text-to-image synthesis, largely attributed to the use of classifier-free guidance (CFG), which enables high-quality, condition-aligned image generation. CFG combines the conditional score (e.g., text-conditioned) with the unconditional score to control the output. However, the unconditional score is in charge of estimating the transition between manifolds of adjacent timesteps from $x_t$ to $x_{t-1}$, which may inadvertently interfere with the trajectory toward the specific condition. In this work, we introduce a novel approach that leverages a geometric perspective on the unconditional score to enhance CFG performance when conditional scores are available. Specifically, we propose a method that filters the singular vectors of both conditional and unconditional scores using singular value decomposition. This filtering process aligns the unconditional score with the conditional score, thereby refining the sampling trajectory to stay closer to the manifold. Our approach improves image quality with negligible additional computation. We provide deeper insights into the score function behavior in diffusion models and present a practical technique for achieving more accurate and contextually coherent image synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TCFG: Tangential Damping Classifier-free Guidance
Kwon, Mingi
Kim, Shin seong
Hsiao, Jaeseok Jeong. Yi Ting
Uh, Youngjung
Computer Vision and Pattern Recognition
Diffusion models have achieved remarkable success in text-to-image synthesis, largely attributed to the use of classifier-free guidance (CFG), which enables high-quality, condition-aligned image generation. CFG combines the conditional score (e.g., text-conditioned) with the unconditional score to control the output. However, the unconditional score is in charge of estimating the transition between manifolds of adjacent timesteps from $x_t$ to $x_{t-1}$, which may inadvertently interfere with the trajectory toward the specific condition. In this work, we introduce a novel approach that leverages a geometric perspective on the unconditional score to enhance CFG performance when conditional scores are available. Specifically, we propose a method that filters the singular vectors of both conditional and unconditional scores using singular value decomposition. This filtering process aligns the unconditional score with the conditional score, thereby refining the sampling trajectory to stay closer to the manifold. Our approach improves image quality with negligible additional computation. We provide deeper insights into the score function behavior in diffusion models and present a practical technique for achieving more accurate and contextually coherent image synthesis.
title TCFG: Tangential Damping Classifier-free Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18137