Infinite-Story: A Training-Free Consistent Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Jihun, Lee, Kyoungmin, Gim, Jongmin, Jo, Hyeonseo, Oh, Minseok, Choi, Wonhyeok, Hwang, Kyumin, Kim, Jaeyeul, Choi, Minwoo, Im, Sunghoon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918205324263424
author Park, Jihun
Lee, Kyoungmin
Gim, Jongmin
Jo, Hyeonseo
Oh, Minseok
Choi, Wonhyeok
Hwang, Kyumin
Kim, Jaeyeul
Choi, Minwoo
Im, Sunghoon
author_facet Park, Jihun
Lee, Kyoungmin
Gim, Jongmin
Jo, Hyeonseo
Oh, Minseok
Choi, Wonhyeok
Hwang, Kyumin
Kim, Jaeyeul
Choi, Minwoo
Im, Sunghoon
contents We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6X faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Infinite-Story: A Training-Free Consistent Text-to-Image Generation
Park, Jihun
Lee, Kyoungmin
Gim, Jongmin
Jo, Hyeonseo
Oh, Minseok
Choi, Wonhyeok
Hwang, Kyumin
Kim, Jaeyeul
Choi, Minwoo
Im, Sunghoon
Computer Vision and Pattern Recognition
We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6X faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.
title Infinite-Story: A Training-Free Consistent Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.13002