GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Kaiyi, Huang, Yukun, Ning, Xuefei, Lin, Zinan, Wang, Yu, Liu, Xihui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910728933343232
author Huang, Kaiyi
Huang, Yukun
Ning, Xuefei
Lin, Zinan
Wang, Yu
Liu, Xihui
author_facet Huang, Kaiyi
Huang, Yukun
Ning, Xuefei
Lin, Zinan
Wang, Yu
Liu, Xihui
contents Text-to-video generation models have shown significant progress in the recent years. However, they still struggle with generating complex dynamic scenes based on compositional text prompts, such as attribute binding for multiple objects, temporal dynamics associated with different objects, and interactions between objects. Our key motivation is that complex tasks can be decomposed into simpler ones, each handled by a role-specialized MLLM agent. Multiple agents can collaborate together to achieve collective intelligence for complex goals. We propose GenMAC, an iterative, multi-agent framework that enables compositional text-to-video generation. The collaborative workflow includes three stages: Design, Generation, and Redesign, with an iterative loop between the Generation and Redesign stages to progressively verify and refine the generated videos. The Redesign stage is the most challenging stage that aims to verify the generated videos, suggest corrections, and redesign the text prompts, frame-wise layouts, and guidance scales for the next iteration of generation. To avoid hallucination of a single MLLM agent, we decompose this stage to four sequentially-executed MLLM-based agents: verification agent, suggestion agent, correction agent, and output structuring agent. Furthermore, to tackle diverse scenarios of compositional text-to-video generation, we design a self-routing mechanism to adaptively select the proper correction agent from a collection of correction agents each specialized for one scenario. Extensive experiments demonstrate the effectiveness of GenMAC, achieving state-of-the art performance in compositional text-to-video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04440
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
Huang, Kaiyi
Huang, Yukun
Ning, Xuefei
Lin, Zinan
Wang, Yu
Liu, Xihui
Computer Vision and Pattern Recognition
Text-to-video generation models have shown significant progress in the recent years. However, they still struggle with generating complex dynamic scenes based on compositional text prompts, such as attribute binding for multiple objects, temporal dynamics associated with different objects, and interactions between objects. Our key motivation is that complex tasks can be decomposed into simpler ones, each handled by a role-specialized MLLM agent. Multiple agents can collaborate together to achieve collective intelligence for complex goals. We propose GenMAC, an iterative, multi-agent framework that enables compositional text-to-video generation. The collaborative workflow includes three stages: Design, Generation, and Redesign, with an iterative loop between the Generation and Redesign stages to progressively verify and refine the generated videos. The Redesign stage is the most challenging stage that aims to verify the generated videos, suggest corrections, and redesign the text prompts, frame-wise layouts, and guidance scales for the next iteration of generation. To avoid hallucination of a single MLLM agent, we decompose this stage to four sequentially-executed MLLM-based agents: verification agent, suggestion agent, correction agent, and output structuring agent. Furthermore, to tackle diverse scenarios of compositional text-to-video generation, we design a self-routing mechanism to adaptively select the proper correction agent from a collection of correction agents each specialized for one scenario. Extensive experiments demonstrate the effectiveness of GenMAC, achieving state-of-the art performance in compositional text-to-video generation.
title GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04440