Bridging Your Imagination with Audio-Video Generation via a Unified Director

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Jiaxu, Hu, Tianshu, Zhang, Yuan, Li, Zenan, Luo, Linjie, Lin, Guosheng, Chen, Xin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909977529024512
author Zhang, Jiaxu
Hu, Tianshu
Zhang, Yuan
Li, Zenan
Luo, Linjie
Lin, Guosheng
Chen, Xin
author_facet Zhang, Jiaxu
Hu, Tianshu
Zhang, Yuan
Li, Zenan
Luo, Linjie
Lin, Guosheng
Chen, Xin
contents Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two tasks should be unified within a single framework, as logical reasoning and imaginative thinking are both fundamental qualities of a film director. In this work, we propose UniMAGE, a unified director model that bridges user prompts with well-structured scripts, thereby empowering non-experts to produce long-context, multi-shot films by leveraging existing audio-video generation models. To achieve this, we employ the Mixture-of-Transformers architecture that unifies text and image generation. To further enhance narrative logic and keyframe consistency, we introduce a ``first interleaving, then disentangling'' training paradigm. Specifically, we first perform Interleaved Concept Learning, which utilizes interleaved text-image data to foster the model's deeper understanding and imaginative interpretation of scripts. We then conduct Disentangled Expert Learning, which decouples script writing from keyframe generation, enabling greater flexibility and creativity in storytelling. Extensive experiments demonstrate that UniMAGE achieves state-of-the-art performance among open-source models, generating logically coherent video scripts and visually consistent keyframe images.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging Your Imagination with Audio-Video Generation via a Unified Director
Zhang, Jiaxu
Hu, Tianshu
Zhang, Yuan
Li, Zenan
Luo, Linjie
Lin, Guosheng
Chen, Xin
Computer Vision and Pattern Recognition
Multimedia
Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two tasks should be unified within a single framework, as logical reasoning and imaginative thinking are both fundamental qualities of a film director. In this work, we propose UniMAGE, a unified director model that bridges user prompts with well-structured scripts, thereby empowering non-experts to produce long-context, multi-shot films by leveraging existing audio-video generation models. To achieve this, we employ the Mixture-of-Transformers architecture that unifies text and image generation. To further enhance narrative logic and keyframe consistency, we introduce a ``first interleaving, then disentangling'' training paradigm. Specifically, we first perform Interleaved Concept Learning, which utilizes interleaved text-image data to foster the model's deeper understanding and imaginative interpretation of scripts. We then conduct Disentangled Expert Learning, which decouples script writing from keyframe generation, enabling greater flexibility and creativity in storytelling. Extensive experiments demonstrate that UniMAGE achieves state-of-the-art performance among open-source models, generating logically coherent video scripts and visually consistent keyframe images.
title Bridging Your Imagination with Audio-Video Generation via a Unified Director
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2512.23222