ReNeg: Learning Negative Embedding with Reward Guidance

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Xiaomin, Liu, Yixuan, Isobe, Takashi, Jia, Xu, Cui, Qinpeng, Zhou, Dong, Li, Dong, He, You, Lu, Huchuan, Wang, Zhongdao, Barsoum, Emad
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909655071981568
author Li, Xiaomin
Liu, Yixuan
Isobe, Takashi
Jia, Xu
Cui, Qinpeng
Zhou, Dong
Li, Dong
He, You
Lu, Huchuan
Wang, Zhongdao
Barsoum, Emad
author_facet Li, Xiaomin
Liu, Yixuan
Isobe, Takashi
Jia, Xu
Cui, Qinpeng
Zhou, Dong
Li, Dong
He, You
Lu, Huchuan
Wang, Zhongdao
Barsoum, Emad
contents In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effective learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19637
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReNeg: Learning Negative Embedding with Reward Guidance
Li, Xiaomin
Liu, Yixuan
Isobe, Takashi
Jia, Xu
Cui, Qinpeng
Zhou, Dong
Li, Dong
He, You
Lu, Huchuan
Wang, Zhongdao
Barsoum, Emad
Computer Vision and Pattern Recognition
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effective learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board.
title ReNeg: Learning Negative Embedding with Reward Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19637