SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tu, Rong-Cheng, Sun, Wenhao, Jin, Zhao, Liao, Jingyi, Huang, Jiaxing, Tao, Dacheng
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909407863898112
author Tu, Rong-Cheng
Sun, Wenhao
Jin, Zhao
Liao, Jingyi
Huang, Jiaxing
Tao, Dacheng
author_facet Tu, Rong-Cheng
Sun, Wenhao
Jin, Zhao
Liao, Jingyi
Huang, Jiaxing
Tao, Dacheng
contents While open-source video generation and editing models have made significant progress, individual models are typically limited to specific tasks, failing to meet the diverse needs of users. Effectively coordinating these models can unlock a wide range of video generation and editing capabilities. However, manual coordination is complex and time-consuming, requiring users to deeply understand task requirements and possess comprehensive knowledge of each model's performance, applicability, and limitations, thereby increasing the barrier to entry. To address these challenges, we propose a novel video generation and editing system powered by our Semantic Planning Agent (SPAgent). SPAgent bridges the gap between diverse user intents and the effective utilization of existing generative models, enhancing the adaptability, efficiency, and overall quality of video generation and editing. Specifically, the SPAgent assembles a tool library integrating state-of-the-art open-source image and video generation and editing models as tools. After fine-tuning on our manually annotated dataset, SPAgent can automatically coordinate the tools for video generation and editing, through our novelly designed three-step framework: (1) decoupled intent recognition, (2) principle-guided route planning, and (3) capability-based execution model selection. Additionally, we enhance the SPAgent's video quality evaluation capability, enabling it to autonomously assess and incorporate new video generation and editing models into its tool library without human intervention. Experimental results demonstrate that the SPAgent effectively coordinates models to generate or edit videos, highlighting its versatility and adaptability across various video tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18983
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
Tu, Rong-Cheng
Sun, Wenhao
Jin, Zhao
Liao, Jingyi
Huang, Jiaxing
Tao, Dacheng
Computer Vision and Pattern Recognition
Multiagent Systems
While open-source video generation and editing models have made significant progress, individual models are typically limited to specific tasks, failing to meet the diverse needs of users. Effectively coordinating these models can unlock a wide range of video generation and editing capabilities. However, manual coordination is complex and time-consuming, requiring users to deeply understand task requirements and possess comprehensive knowledge of each model's performance, applicability, and limitations, thereby increasing the barrier to entry. To address these challenges, we propose a novel video generation and editing system powered by our Semantic Planning Agent (SPAgent). SPAgent bridges the gap between diverse user intents and the effective utilization of existing generative models, enhancing the adaptability, efficiency, and overall quality of video generation and editing. Specifically, the SPAgent assembles a tool library integrating state-of-the-art open-source image and video generation and editing models as tools. After fine-tuning on our manually annotated dataset, SPAgent can automatically coordinate the tools for video generation and editing, through our novelly designed three-step framework: (1) decoupled intent recognition, (2) principle-guided route planning, and (3) capability-based execution model selection. Additionally, we enhance the SPAgent's video quality evaluation capability, enabling it to autonomously assess and incorporate new video generation and editing models into its tool library without human intervention. Experimental results demonstrate that the SPAgent effectively coordinates models to generate or edit videos, highlighting its versatility and adaptability across various video tasks.
title SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
topic Computer Vision and Pattern Recognition
Multiagent Systems
url https://arxiv.org/abs/2411.18983