SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917206997073920 |
|---|---|
| author | Grimal, Paul Soumm, Michaël Borgne, Hervé Le Ferret, Olivier Sugimoto, Akihiro |
| author_facet | Grimal, Paul Soumm, Michaël Borgne, Hervé Le Ferret, Olivier Sugimoto, Akihiro |
| contents | State-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel approach that learns a high-success-rate distribution conditioned on a target prompt, ensuring that generated images faithfully reflect the corresponding prompts. Our method explicitly models the signal component during the denoising process, offering fine-grained control that mitigates over-optimization and out-of-distribution artifacts. Moreover, our framework is training-free and seamlessly integrates with both existing diffusion and flow matching architectures. It also supports additional conditioning modalities -- such as bounding boxes -- for enhanced spatial alignment. Extensive experiments demonstrate that our approach outperforms current state-of-the-art methods. The code is available at https://github.com/grimalPaul/gsn-factory. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_13866 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation Grimal, Paul Soumm, Michaël Borgne, Hervé Le Ferret, Olivier Sugimoto, Akihiro Computer Vision and Pattern Recognition State-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel approach that learns a high-success-rate distribution conditioned on a target prompt, ensuring that generated images faithfully reflect the corresponding prompts. Our method explicitly models the signal component during the denoising process, offering fine-grained control that mitigates over-optimization and out-of-distribution artifacts. Moreover, our framework is training-free and seamlessly integrates with both existing diffusion and flow matching architectures. It also supports additional conditioning modalities -- such as bounding boxes -- for enhanced spatial alignment. Extensive experiments demonstrate that our approach outperforms current state-of-the-art methods. The code is available at https://github.com/grimalPaul/gsn-factory. |
| title | SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.13866 |