DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918235562049536 |
|---|---|
| author | Song, Jaewoo Choi, Jooyoung Baek, Kanghyun Lee, Sangyub Park, Daemin Yoon, Sungroh |
| author_facet | Song, Jaewoo Choi, Jooyoung Baek, Kanghyun Lee, Sangyub Park, Daemin Yoon, Sungroh |
| contents | Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a training-free visual text generation method that adopts a divide-and-conquer strategy, leveraging the reliable short-text generation of Multi-Modal Diffusion Transformers. Our method first decomposes a prompt by extracting and dividing the target text, then assigns each to a designated region. To accurately render each segment within their regions while preserving overall image coherence, we introduce two attention masks - Text-Focus and Context-Expansion - applied sequentially during denoising. Additionally, Localized Noise Initialization further improves text accuracy and region alignment without increasing computational cost. Extensive experiments on single- and multisentence benchmarks show that DCText achieves the best text accuracy without compromising image quality while also delivering the lowest generation latency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_01302 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy Song, Jaewoo Choi, Jooyoung Baek, Kanghyun Lee, Sangyub Park, Daemin Yoon, Sungroh Computer Vision and Pattern Recognition Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a training-free visual text generation method that adopts a divide-and-conquer strategy, leveraging the reliable short-text generation of Multi-Modal Diffusion Transformers. Our method first decomposes a prompt by extracting and dividing the target text, then assigns each to a designated region. To accurately render each segment within their regions while preserving overall image coherence, we introduce two attention masks - Text-Focus and Context-Expansion - applied sequentially during denoising. Additionally, Localized Noise Initialization further improves text accuracy and region alignment without increasing computational cost. Extensive experiments on single- and multisentence benchmarks show that DCText achieves the best text accuracy without compromising image quality while also delivering the lowest generation latency. |
| title | DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.01302 |