LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gani, Hanan, Bhat, Shariq Farooq, Naseer, Muzammal, Khan, Salman, Wonka, Peter
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911783399194624
author Gani, Hanan
Bhat, Shariq Farooq
Naseer, Muzammal
Khan, Salman
Wonka, Peter
author_facet Gani, Hanan
Bhat, Shariq Farooq
Naseer, Muzammal
Khan, Salman
Wonka, Peter
contents Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these models often struggle to faithfully capture all the nuanced details within longer and more elaborate textual inputs. In response, we present a novel approach leveraging Large Language Models (LLMs) to extract critical components from text prompts, including bounding box coordinates for foreground objects, detailed textual descriptions for individual objects, and a succinct background context. These components form the foundation of our layout-to-image generation model, which operates in two phases. The initial Global Scene Generation utilizes object layouts and background context to create an initial scene but often falls short in faithfully representing object characteristics as specified in the prompts. To address this limitation, we introduce an Iterative Refinement Scheme that iteratively evaluates and refines box-level content to align them with their textual descriptions, recomposing objects as needed to ensure consistency. Our evaluation on complex prompts featuring multiple objects demonstrates a substantial improvement in recall compared to baseline diffusion models. This is further validated by a user study, underscoring the efficacy of our approach in generating coherent and detailed scenes from intricate textual inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10640
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
Gani, Hanan
Bhat, Shariq Farooq
Naseer, Muzammal
Khan, Salman
Wonka, Peter
Computer Vision and Pattern Recognition
Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these models often struggle to faithfully capture all the nuanced details within longer and more elaborate textual inputs. In response, we present a novel approach leveraging Large Language Models (LLMs) to extract critical components from text prompts, including bounding box coordinates for foreground objects, detailed textual descriptions for individual objects, and a succinct background context. These components form the foundation of our layout-to-image generation model, which operates in two phases. The initial Global Scene Generation utilizes object layouts and background context to create an initial scene but often falls short in faithfully representing object characteristics as specified in the prompts. To address this limitation, we introduce an Iterative Refinement Scheme that iteratively evaluates and refines box-level content to align them with their textual descriptions, recomposing objects as needed to ensure consistency. Our evaluation on complex prompts featuring multiple objects demonstrates a substantial improvement in recall compared to baseline diffusion models. This is further validated by a user study, underscoring the efficacy of our approach in generating coherent and detailed scenes from intricate textual inputs.
title LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.10640