Infinite-Story: A Training-Free Consistent Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918205324263424 |
|---|---|
| author | Park, Jihun Lee, Kyoungmin Gim, Jongmin Jo, Hyeonseo Oh, Minseok Choi, Wonhyeok Hwang, Kyumin Kim, Jaeyeul Choi, Minwoo Im, Sunghoon |
| author_facet | Park, Jihun Lee, Kyoungmin Gim, Jongmin Jo, Hyeonseo Oh, Minseok Choi, Wonhyeok Hwang, Kyumin Kim, Jaeyeul Choi, Minwoo Im, Sunghoon |
| contents | We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6X faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_13002 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Infinite-Story: A Training-Free Consistent Text-to-Image Generation Park, Jihun Lee, Kyoungmin Gim, Jongmin Jo, Hyeonseo Oh, Minseok Choi, Wonhyeok Hwang, Kyumin Kim, Jaeyeul Choi, Minwoo Im, Sunghoon Computer Vision and Pattern Recognition We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6X faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling. |
| title | Infinite-Story: A Training-Free Consistent Text-to-Image Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.13002 |