Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Akdemir, Kiymet, Kazimi, Tahira, Yanardag, Pinar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912445129293824
author Akdemir, Kiymet
Kazimi, Tahira
Yanardag, Pinar
author_facet Akdemir, Kiymet
Kazimi, Tahira
Yanardag, Pinar
contents Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects persist and evolve throughout the story. Despite recent advances in diffusion models, current approaches often fail to preserve key character attributes, leading to incoherent narratives. In this work, we propose a collaborative multi-agent framework that autonomously identifies, corrects, and refines inconsistencies across multi-panel story visualizations. The agents operate in an iterative loop, enabling fine-grained, panel-level updates without re-generating entire sequences. Our framework is model-agnostic and flexibly integrates with a variety of diffusion models, including rectified flow transformers such as Flux and latent diffusion models such as Stable Diffusion. Quantitative and qualitative experiments show that our method outperforms prior approaches in terms of multi-panel consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models
Akdemir, Kiymet
Kazimi, Tahira
Yanardag, Pinar
Computer Vision and Pattern Recognition
Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects persist and evolve throughout the story. Despite recent advances in diffusion models, current approaches often fail to preserve key character attributes, leading to incoherent narratives. In this work, we propose a collaborative multi-agent framework that autonomously identifies, corrects, and refines inconsistencies across multi-panel story visualizations. The agents operate in an iterative loop, enabling fine-grained, panel-level updates without re-generating entire sequences. Our framework is model-agnostic and flexibly integrates with a variety of diffusion models, including rectified flow transformers such as Flux and latent diffusion models such as Stable Diffusion. Quantitative and qualitative experiments show that our method outperforms prior approaches in terms of multi-panel consistency.
title Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18900