The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kodathala, Sai Varun, Vunnam, Rakesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915506930319360
author Kodathala, Sai Varun
Vunnam, Rakesh
author_facet Kodathala, Sai Varun
Vunnam, Rakesh
contents With the increasing integration of multimodal AI systems in creative workflows, understanding information loss in vision-language-vision pipelines has become important for evaluating system limitations. However, the degradation that occurs when visual content passes through textual intermediation remains poorly quantified. In this work, we provide empirical analysis of the describe-then-generate bottleneck, where natural language serves as an intermediate representation for visual information. We generated 150 image pairs through the describe-then-generate pipeline and applied existing metrics (LPIPS, SSIM, and color distance) to measure information preservation across perceptual, structural, and chromatic dimensions. Our evaluation reveals that 99.3% of samples exhibit substantial perceptual degradation and 91.5% demonstrate significant structural information loss, providing empirical evidence that the describe-then-generate bottleneck represents a measurable and consistent limitation in contemporary multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
Kodathala, Sai Varun
Vunnam, Rakesh
Computer Vision and Pattern Recognition
Artificial Intelligence
With the increasing integration of multimodal AI systems in creative workflows, understanding information loss in vision-language-vision pipelines has become important for evaluating system limitations. However, the degradation that occurs when visual content passes through textual intermediation remains poorly quantified. In this work, we provide empirical analysis of the describe-then-generate bottleneck, where natural language serves as an intermediate representation for visual information. We generated 150 image pairs through the describe-then-generate pipeline and applied existing metrics (LPIPS, SSIM, and color distance) to measure information preservation across perceptual, structural, and chromatic dimensions. Our evaluation reveals that 99.3% of samples exhibit substantial perceptual degradation and 91.5% demonstrate significant structural information loss, providing empirical evidence that the describe-then-generate bottleneck represents a measurable and consistent limitation in contemporary multimodal systems.
title The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.18179