StoryMem: Multi-shot Long Video Storytelling with Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaiwen, Jiang, Liming, Wang, Angtian, Fang, Jacob Zhiyuan, Zhi, Tiancheng, Yan, Qing, Kang, Hao, Lu, Xin, Pan, Xingang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918259425542144
author Zhang, Kaiwen
Jiang, Liming
Wang, Angtian
Fang, Jacob Zhiyuan
Zhi, Tiancheng
Yan, Qing
Kang, Hao
Lu, Xin
Pan, Xingang
author_facet Zhang, Kaiwen
Jiang, Liming
Wang, Angtian
Fang, Jacob Zhiyuan
Zhi, Tiancheng
Yan, Qing
Kang, Hao
Lu, Xin
Pan, Xingang
contents Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot synthesis conditioned on explicit visual memory, transforming pre-trained single-shot video diffusion models into multi-shot storytellers. This is achieved by a novel Memory-to-Video (M2V) design, which maintains a compact and dynamically updated memory bank of keyframes from historical generated shots. The stored memory is then injected into single-shot video diffusion models via latent concatenation and negative RoPE shifts with only LoRA fine-tuning. A semantic keyframe selection strategy, together with aesthetic preference filtering, further ensures informative and stable memory throughout generation. Moreover, the proposed framework naturally accommodates smooth shot transitions and customized story generation applications. To facilitate evaluation, we introduce ST-Bench, a diverse benchmark for multi-shot video storytelling. Extensive experiments demonstrate that StoryMem achieves superior cross-shot consistency over previous methods while preserving high aesthetic quality and prompt adherence, marking a significant step toward coherent minute-long video storytelling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StoryMem: Multi-shot Long Video Storytelling with Memory
Zhang, Kaiwen
Jiang, Liming
Wang, Angtian
Fang, Jacob Zhiyuan
Zhi, Tiancheng
Yan, Qing
Kang, Hao
Lu, Xin
Pan, Xingang
Computer Vision and Pattern Recognition
Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot synthesis conditioned on explicit visual memory, transforming pre-trained single-shot video diffusion models into multi-shot storytellers. This is achieved by a novel Memory-to-Video (M2V) design, which maintains a compact and dynamically updated memory bank of keyframes from historical generated shots. The stored memory is then injected into single-shot video diffusion models via latent concatenation and negative RoPE shifts with only LoRA fine-tuning. A semantic keyframe selection strategy, together with aesthetic preference filtering, further ensures informative and stable memory throughout generation. Moreover, the proposed framework naturally accommodates smooth shot transitions and customized story generation applications. To facilitate evaluation, we introduce ST-Bench, a diverse benchmark for multi-shot video storytelling. Extensive experiments demonstrate that StoryMem achieves superior cross-shot consistency over previous methods while preserving high aesthetic quality and prompt adherence, marking a significant step toward coherent minute-long video storytelling.
title StoryMem: Multi-shot Long Video Storytelling with Memory
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19539