CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Qinglin, Cai, Kaitong, Chen, Ruiqi, Lv, Qinhan, Wang, Keze
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911341018611712
author Zeng, Qinglin
Cai, Kaitong
Chen, Ruiqi
Lv, Qinhan
Wang, Keze
author_facet Zeng, Qinglin
Cai, Kaitong
Chen, Ruiqi
Lv, Qinhan
Wang, Keze
contents Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and unstable temporal structure. We propose CoAgent, a collaborative and closed-loop framework for coherent video generation that formulates the process as a plan-synthesize-verify pipeline. Given a user prompt, style reference, and pacing constraints, a Storyboard Planner decomposes the input into structured shot-level plans with explicit entities, spatial relations, and temporal cues. A Global Context Manager maintains entity-level memory to preserve appearance and identity consistency across shots. Each shot is then generated by a Synthesis Module under the guidance of a Visual Consistency Controller, while a Verifier Agent evaluates intermediate results using vision-language reasoning and triggers selective regeneration when inconsistencies are detected. Finally, a pacing-aware editor refines temporal rhythm and transitions to match the desired narrative flow. Extensive experiments demonstrate that CoAgent significantly improves coherence, visual consistency, and narrative quality in long-form video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22536
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
Zeng, Qinglin
Cai, Kaitong
Chen, Ruiqi
Lv, Qinhan
Wang, Keze
Computer Vision and Pattern Recognition
Artificial Intelligence
Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and unstable temporal structure. We propose CoAgent, a collaborative and closed-loop framework for coherent video generation that formulates the process as a plan-synthesize-verify pipeline. Given a user prompt, style reference, and pacing constraints, a Storyboard Planner decomposes the input into structured shot-level plans with explicit entities, spatial relations, and temporal cues. A Global Context Manager maintains entity-level memory to preserve appearance and identity consistency across shots. Each shot is then generated by a Synthesis Module under the guidance of a Visual Consistency Controller, while a Verifier Agent evaluates intermediate results using vision-language reasoning and triggers selective regeneration when inconsistencies are detected. Finally, a pacing-aware editor refines temporal rhythm and transitions to match the desired narrative flow. Extensive experiments demonstrate that CoAgent significantly improves coherence, visual consistency, and narrative quality in long-form video generation.
title CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.22536