GPTDrawer: Enhancing Visual Synthesis through ChatGPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Kun, Chen, Xinwei, Song, Tianyou, Zhang, Hansong, Zhang, Wenzhe, Shan, Qing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916523202838528
author Li, Kun
Chen, Xinwei
Song, Tianyou
Zhang, Hansong
Zhang, Wenzhe
Shan, Qing
author_facet Li, Kun
Chen, Xinwei
Song, Tianyou
Zhang, Hansong
Zhang, Wenzhe
Shan, Qing
contents In the burgeoning field of AI-driven image generation, the quest for precision and relevance in response to textual prompts remains paramount. This paper introduces GPTDrawer, an innovative pipeline that leverages the generative prowess of GPT-based models to enhance the visual synthesis process. Our methodology employs a novel algorithm that iteratively refines input prompts using keyword extraction, semantic analysis, and image-text congruence evaluation. By integrating ChatGPT for natural language processing and Stable Diffusion for image generation, GPTDrawer produces a batch of images that undergo successive refinement cycles, guided by cosine similarity metrics until a threshold of semantic alignment is attained. The results demonstrate a marked improvement in the fidelity of images generated in accordance with user-defined prompts, showcasing the system's ability to interpret and visualize complex semantic constructs. The implications of this work extend to various applications, from creative arts to design automation, setting a new benchmark for AI-assisted creative processes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10429
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GPTDrawer: Enhancing Visual Synthesis through ChatGPT
Li, Kun
Chen, Xinwei
Song, Tianyou
Zhang, Hansong
Zhang, Wenzhe
Shan, Qing
Computer Vision and Pattern Recognition
Artificial Intelligence
In the burgeoning field of AI-driven image generation, the quest for precision and relevance in response to textual prompts remains paramount. This paper introduces GPTDrawer, an innovative pipeline that leverages the generative prowess of GPT-based models to enhance the visual synthesis process. Our methodology employs a novel algorithm that iteratively refines input prompts using keyword extraction, semantic analysis, and image-text congruence evaluation. By integrating ChatGPT for natural language processing and Stable Diffusion for image generation, GPTDrawer produces a batch of images that undergo successive refinement cycles, guided by cosine similarity metrics until a threshold of semantic alignment is attained. The results demonstrate a marked improvement in the fidelity of images generated in accordance with user-defined prompts, showcasing the system's ability to interpret and visualize complex semantic constructs. The implications of this work extend to various applications, from creative arts to design automation, setting a new benchmark for AI-assisted creative processes.
title GPTDrawer: Enhancing Visual Synthesis through ChatGPT
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.10429