ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910629030264832 |
|---|---|
| author | Gal, Rinon Haviv, Adi Alaluf, Yuval Bermano, Amit H. Cohen-Or, Daniel Chechik, Gal |
| author_facet | Gal, Rinon Haviv, Adi Alaluf, Yuval Bermano, Amit H. Cohen-Or, Daniel Chechik, Gal |
| contents | The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_01731 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation Gal, Rinon Haviv, Adi Alaluf, Yuval Bermano, Amit H. Cohen-Or, Daniel Chechik, Gal Computer Vision and Pattern Recognition Computation and Language Graphics The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field. |
| title | ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation |
| topic | Computer Vision and Pattern Recognition Computation and Language Graphics |
| url | https://arxiv.org/abs/2410.01731 |