CustomText: Customized Textual Image Generation using Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Paliwal, Shubham, Jain, Arushi, Sharma, Monika, Jamwal, Vikram, Vig, Lovekesh
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914803891568640
author Paliwal, Shubham
Jain, Arushi
Sharma, Monika
Jamwal, Vikram
Vig, Lovekesh
author_facet Paliwal, Shubham
Jain, Arushi
Sharma, Monika
Jamwal, Vikram
Vig, Lovekesh
contents Textual image generation spans diverse fields like advertising, education, product packaging, social media, information visualization, and branding. Despite recent strides in language-guided image synthesis using diffusion models, current models excel in image generation but struggle with accurate text rendering and offer limited control over font attributes. In this paper, we aim to enhance the synthesis of high-quality images with precise text customization, thereby contributing to the advancement of image generation models. We call our proposed method CustomText. Our implementation leverages a pre-trained TextDiffuser model to enable control over font color, background, and types. Additionally, to address the challenge of accurately rendering small-sized fonts, we train the ControlNet model for a consistency decoder, significantly enhancing text-generation performance. We assess the performance of CustomText in comparison to previous methods of textual image generation on the publicly available CTW-1500 dataset and a self-curated dataset for small-text generation, showcasing superior results.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12531
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CustomText: Customized Textual Image Generation using Diffusion Models
Paliwal, Shubham
Jain, Arushi
Sharma, Monika
Jamwal, Vikram
Vig, Lovekesh
Computer Vision and Pattern Recognition
Machine Learning
Textual image generation spans diverse fields like advertising, education, product packaging, social media, information visualization, and branding. Despite recent strides in language-guided image synthesis using diffusion models, current models excel in image generation but struggle with accurate text rendering and offer limited control over font attributes. In this paper, we aim to enhance the synthesis of high-quality images with precise text customization, thereby contributing to the advancement of image generation models. We call our proposed method CustomText. Our implementation leverages a pre-trained TextDiffuser model to enable control over font color, background, and types. Additionally, to address the challenge of accurately rendering small-sized fonts, we train the ControlNet model for a consistency decoder, significantly enhancing text-generation performance. We assess the performance of CustomText in comparison to previous methods of textual image generation on the publicly available CTW-1500 dataset and a self-curated dataset for small-text generation, showcasing superior results.
title CustomText: Customized Textual Image Generation using Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.12531