DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Jaewoo, Choi, Jooyoung, Baek, Kanghyun, Lee, Sangyub, Park, Daemin, Yoon, Sungroh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918235562049536
author Song, Jaewoo
Choi, Jooyoung
Baek, Kanghyun
Lee, Sangyub
Park, Daemin
Yoon, Sungroh
author_facet Song, Jaewoo
Choi, Jooyoung
Baek, Kanghyun
Lee, Sangyub
Park, Daemin
Yoon, Sungroh
contents Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a training-free visual text generation method that adopts a divide-and-conquer strategy, leveraging the reliable short-text generation of Multi-Modal Diffusion Transformers. Our method first decomposes a prompt by extracting and dividing the target text, then assigns each to a designated region. To accurately render each segment within their regions while preserving overall image coherence, we introduce two attention masks - Text-Focus and Context-Expansion - applied sequentially during denoising. Additionally, Localized Noise Initialization further improves text accuracy and region alignment without increasing computational cost. Extensive experiments on single- and multisentence benchmarks show that DCText achieves the best text accuracy without compromising image quality while also delivering the lowest generation latency.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
Song, Jaewoo
Choi, Jooyoung
Baek, Kanghyun
Lee, Sangyub
Park, Daemin
Yoon, Sungroh
Computer Vision and Pattern Recognition
Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a training-free visual text generation method that adopts a divide-and-conquer strategy, leveraging the reliable short-text generation of Multi-Modal Diffusion Transformers. Our method first decomposes a prompt by extracting and dividing the target text, then assigns each to a designated region. To accurately render each segment within their regions while preserving overall image coherence, we introduce two attention masks - Text-Focus and Context-Expansion - applied sequentially during denoising. Additionally, Localized Noise Initialization further improves text accuracy and region alignment without increasing computational cost. Extensive experiments on single- and multisentence benchmarks show that DCText achieves the best text accuracy without compromising image quality while also delivering the lowest generation latency.
title DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01302