Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jinzhuo, Zhang, Jiangning, Jiang, Wencan, Wang, Yabiao, Liang, Dingkang, Xue, Zhucun, Yi, Ran, Liu, Yong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914578598723584
author Liu, Jinzhuo
Zhang, Jiangning
Jiang, Wencan
Wang, Yabiao
Liang, Dingkang
Xue, Zhucun
Yi, Ran
Liu, Yong
author_facet Liu, Jinzhuo
Zhang, Jiangning
Jiang, Wencan
Wang, Yabiao
Liang, Dingkang
Xue, Zhucun
Yi, Ran
Liu, Yong
contents Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined strategies or retrieve keyframes based on coarse implicit attention signals, both of which fail to handle evolving prompts with shifting entity references, leading to identity drift, character duplication, and attribute loss. To address this, we propose IAMFlow, a training-free identity-aware memory framework that explicitly models and tracks persistent entity identities, enabling consistent generation across prompt transitions. Specifically, an LLM extracts entities with visual attributes from each prompt and assigns unique global IDs for identity-aware memory, while a VLM asynchronously verifies and refines attributes from rendered frames, enabling explicit entity tracking in place of implicit similarity-based matching. To keep the proposed framework computationally practical, we design a systematic inference acceleration pipeline, including asynchronous visual verification, adaptive prompt transition, and model quantization, which achieves faster generation than existing baselines. Furthermore, we introduce NarraStream-Bench, a benchmark for narrative streaming video generation that features 324 multi-prompt scripts spanning six dimensions and a three-dimensional evaluation protocol that integrates both traditional metrics and multimodal large language model-based assessments. Extensive experiments show that IAMFlow, despite being training-free, achieves the best overall performance on NarraStream-Bench, outperforming the strongest baseline by 2.56 points, while achieving a 1.39$\times$ speedup over the most efficient baseline in the 60-second multi-prompt setting.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
Liu, Jinzhuo
Zhang, Jiangning
Jiang, Wencan
Wang, Yabiao
Liang, Dingkang
Xue, Zhucun
Yi, Ran
Liu, Yong
Computer Vision and Pattern Recognition
Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined strategies or retrieve keyframes based on coarse implicit attention signals, both of which fail to handle evolving prompts with shifting entity references, leading to identity drift, character duplication, and attribute loss. To address this, we propose IAMFlow, a training-free identity-aware memory framework that explicitly models and tracks persistent entity identities, enabling consistent generation across prompt transitions. Specifically, an LLM extracts entities with visual attributes from each prompt and assigns unique global IDs for identity-aware memory, while a VLM asynchronously verifies and refines attributes from rendered frames, enabling explicit entity tracking in place of implicit similarity-based matching. To keep the proposed framework computationally practical, we design a systematic inference acceleration pipeline, including asynchronous visual verification, adaptive prompt transition, and model quantization, which achieves faster generation than existing baselines. Furthermore, we introduce NarraStream-Bench, a benchmark for narrative streaming video generation that features 324 multi-prompt scripts spanning six dimensions and a three-dimensional evaluation protocol that integrates both traditional metrics and multimodal large language model-based assessments. Extensive experiments show that IAMFlow, despite being training-free, achieves the best overall performance on NarraStream-Bench, outperforming the strongest baseline by 2.56 points, while achieving a 1.39$\times$ speedup over the most efficient baseline in the 60-second multi-prompt setting.
title Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18733