Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lv, Zheqi, Chen, Junhao, Tian, Qi, Yin, Keting, Zhang, Shengyu, Wu, Fei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908380221669376
author Lv, Zheqi
Chen, Junhao
Tian, Qi
Yin, Keting
Zhang, Shengyu
Wu, Fei
author_facet Lv, Zheqi
Chen, Junhao
Tian, Qi
Yin, Keting
Zhang, Shengyu
Wu, Fei
contents Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current inference pipelines generally lack interpretable semantic supervision and correction mechanisms throughout the denoising process. Most existing approaches rely solely on post-hoc scoring of the final image, prompt filtering, or heuristic resampling strategies-making them ineffective in providing actionable guidance for correcting the generative trajectory. As a result, models often suffer from object confusion, spatial errors, inaccurate counts, and missing semantic elements, severely compromising prompt-image alignment and image quality. To tackle these challenges, we propose MLLM Semantic-Corrected Ping-Pong-Ahead Diffusion (PPAD), a novel framework that, for the first time, introduces a Multimodal Large Language Model (MLLM) as a semantic observer during inference. PPAD performs real-time analysis on intermediate generations, identifies latent semantic inconsistencies, and translates feedback into controllable signals that actively guide the remaining denoising steps. The framework supports both inference-only and training-enhanced settings, and performs semantic correction at only extremely few diffusion steps, offering strong generality and scalability. Extensive experiments demonstrate PPAD's significant improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20053
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
Lv, Zheqi
Chen, Junhao
Tian, Qi
Yin, Keting
Zhang, Shengyu
Wu, Fei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current inference pipelines generally lack interpretable semantic supervision and correction mechanisms throughout the denoising process. Most existing approaches rely solely on post-hoc scoring of the final image, prompt filtering, or heuristic resampling strategies-making them ineffective in providing actionable guidance for correcting the generative trajectory. As a result, models often suffer from object confusion, spatial errors, inaccurate counts, and missing semantic elements, severely compromising prompt-image alignment and image quality. To tackle these challenges, we propose MLLM Semantic-Corrected Ping-Pong-Ahead Diffusion (PPAD), a novel framework that, for the first time, introduces a Multimodal Large Language Model (MLLM) as a semantic observer during inference. PPAD performs real-time analysis on intermediate generations, identifies latent semantic inconsistencies, and translates feedback into controllable signals that actively guide the remaining denoising steps. The framework supports both inference-only and training-enhanced settings, and performs semantic correction at only extremely few diffusion steps, offering strong generality and scalability. Extensive experiments demonstrate PPAD's significant improvements.
title Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2505.20053