AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yan, Hailong, Liu, Shice, Wang, Tao, Zhang, Xiangtao, Zhong, Yijie, Chen, Jinwei, Zhang, Le, Li, Bo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917291250155520
author Yan, Hailong
Liu, Shice
Wang, Tao
Zhang, Xiangtao
Zhong, Yijie
Chen, Jinwei
Zhang, Le
Li, Bo
author_facet Yan, Hailong
Liu, Shice
Wang, Tao
Zhang, Xiangtao
Zhong, Yijie
Chen, Jinwei
Zhang, Le
Li, Bo
contents Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three key limitations: (1) Static models lack dynamic expressiveness and often resort to "copy-paste" pattern. (2) One-shot inference cannot iteratively correct missing attributes or poor prompt adherence. (3) Multi-agents rely on non-robust evaluators, ill-suited for assessing stylized, non-realistic animation. To address these, we propose AnimeAgent, the first Image-to-Video (I2V)-based multi-agent framework for CSG. Inspired by Disney's "Combination of Straight Ahead and Pose to Pose" workflow, AnimeAgent leverages I2V's implicit motion prior to enhance consistency and expressiveness, while a mixed subjective-objective reviewer enables reliable iterative refinement. We also collect a human-annotated CSG benchmark with ground-truth. Experiments show AnimeAgent achieves SOTA performance in consistency, prompt fidelity, and stylization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20664
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist?
Yan, Hailong
Liu, Shice
Wang, Tao
Zhang, Xiangtao
Zhong, Yijie
Chen, Jinwei
Zhang, Le
Li, Bo
Computer Vision and Pattern Recognition
Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three key limitations: (1) Static models lack dynamic expressiveness and often resort to "copy-paste" pattern. (2) One-shot inference cannot iteratively correct missing attributes or poor prompt adherence. (3) Multi-agents rely on non-robust evaluators, ill-suited for assessing stylized, non-realistic animation. To address these, we propose AnimeAgent, the first Image-to-Video (I2V)-based multi-agent framework for CSG. Inspired by Disney's "Combination of Straight Ahead and Pose to Pose" workflow, AnimeAgent leverages I2V's implicit motion prior to enhance consistency and expressiveness, while a mixed subjective-objective reviewer enables reliable iterative refinement. We also collect a human-annotated CSG benchmark with ground-truth. Experiments show AnimeAgent achieves SOTA performance in consistency, prompt fidelity, and stylization.
title AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.20664