ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gal, Rinon, Haviv, Adi, Alaluf, Yuval, Bermano, Amit H., Cohen-Or, Daniel, Chechik, Gal
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910629030264832
author Gal, Rinon
Haviv, Adi
Alaluf, Yuval
Bermano, Amit H.
Cohen-Or, Daniel
Chechik, Gal
author_facet Gal, Rinon
Haviv, Adi
Alaluf, Yuval
Bermano, Amit H.
Cohen-Or, Daniel
Chechik, Gal
contents The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01731
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
Gal, Rinon
Haviv, Adi
Alaluf, Yuval
Bermano, Amit H.
Cohen-Or, Daniel
Chechik, Gal
Computer Vision and Pattern Recognition
Computation and Language
Graphics
The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field.
title ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Computation and Language
Graphics
url https://arxiv.org/abs/2410.01731