Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.19637 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909655071981568 |
|---|---|
| author | Li, Xiaomin Liu, Yixuan Isobe, Takashi Jia, Xu Cui, Qinpeng Zhou, Dong Li, Dong He, You Lu, Huchuan Wang, Zhongdao Barsoum, Emad |
| author_facet | Li, Xiaomin Liu, Yixuan Isobe, Takashi Jia, Xu Cui, Qinpeng Zhou, Dong Li, Dong He, You Lu, Huchuan Wang, Zhongdao Barsoum, Emad |
| contents | In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effective learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_19637 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ReNeg: Learning Negative Embedding with Reward Guidance Li, Xiaomin Liu, Yixuan Isobe, Takashi Jia, Xu Cui, Qinpeng Zhou, Dong Li, Dong He, You Lu, Huchuan Wang, Zhongdao Barsoum, Emad Computer Vision and Pattern Recognition In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effective learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. |
| title | ReNeg: Learning Negative Embedding with Reward Guidance |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.19637 |