Saved in:
Bibliographic Details
Main Authors: Yao, Yuxuan, Chen, Yuxuan, Li, Hui, Cheng, Kaihui, Guo, Qipeng, Sun, Yuwei, Dong, Zilong, Wang, Jingdong, Zhu, Siyu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.06886
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt representation in the text branch is progressively forgotten as depth increases. We further verify this effect on three representative MMDiTs--SD3, SD3.5, and FLUX.1 by probing linguistic attributes of the representations over the layers in the text branch. Motivated by these findings, we introduce a training-free approach, prompt reinjection, which reinjects prompt representations from early layers into later layers to alleviate this forgetting. Experiments on GenEval, DPG, and T2I-CompBench++ show consistent gains in instruction-following capability, along with improvements on metrics capturing preference, aesthetics, and overall text--image generation quality.