SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912671900631040 |
|---|---|
| author | Song, Quanjian Zhou, Donghao Lin, Jingyu Shen, Fei Wang, Jiaze Hu, Xiaowei Chen, Cunjian Heng, Pheng-Ann |
| author_facet | Song, Quanjian Zhou, Donghao Lin, Jingyu Shen, Fei Wang, Jiaze Hu, Xiaowei Chen, Cunjian Heng, Pheng-Ann |
| contents | Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22994 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency Song, Quanjian Zhou, Donghao Lin, Jingyu Shen, Fei Wang, Jiaze Hu, Xiaowei Chen, Cunjian Heng, Pheng-Ann Computer Vision and Pattern Recognition Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games. |
| title | SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.22994 |