A Versatile Multimodal Agent for Multimedia Content Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Daoan, Yao, Wenlin, Wang, Xiaoyang, Hu, Yebowen, Luo, Jiebo, Yu, Dong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909982759321600
author Zhang, Daoan
Yao, Wenlin
Wang, Xiaoyang
Hu, Yebowen
Luo, Jiebo
Yu, Dong
author_facet Zhang, Daoan
Yao, Wenlin
Wang, Xiaoyang
Hu, Yebowen
Luo, Jiebo
Yu, Dong
contents With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of current AIGC models, most models can only serve as individual components within specific application scenarios and are not capable of completing tasks end-to-end in real-world applications. In real-world applications, editing experts often work with a wide variety of images and video inputs, producing multimodal outputs -- a video typically includes audio, text, and other elements. This level of integration across multiple modalities is something current models are unable to achieve effectively. However, the rise of agent-based systems has made it possible to use AI tools to tackle complex content generation tasks. To deal with the complex scenarios, in this paper, we propose a MultiMedia-Agent designed to automate complex content creation. Our agent system includes a data generation pipeline, a tool library for content creation, and a set of metrics for evaluating preference alignment. Notably, we introduce the skill acquisition theory to model the training data curation and agent training. We designed a two-stage correlation strategy for plan optimization, including self-correlation and model preference correlation. Additionally, we utilized the generated plans to train the MultiMedia-Agent via a three stage approach including base/success plan finetune and preference optimization. The comparison results demonstrate that the our approaches are effective and the MultiMedia-Agent can generate better multimedia content compared to novel models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03250
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Versatile Multimodal Agent for Multimedia Content Generation
Zhang, Daoan
Yao, Wenlin
Wang, Xiaoyang
Hu, Yebowen
Luo, Jiebo
Yu, Dong
Computer Vision and Pattern Recognition
With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of current AIGC models, most models can only serve as individual components within specific application scenarios and are not capable of completing tasks end-to-end in real-world applications. In real-world applications, editing experts often work with a wide variety of images and video inputs, producing multimodal outputs -- a video typically includes audio, text, and other elements. This level of integration across multiple modalities is something current models are unable to achieve effectively. However, the rise of agent-based systems has made it possible to use AI tools to tackle complex content generation tasks. To deal with the complex scenarios, in this paper, we propose a MultiMedia-Agent designed to automate complex content creation. Our agent system includes a data generation pipeline, a tool library for content creation, and a set of metrics for evaluating preference alignment. Notably, we introduce the skill acquisition theory to model the training data curation and agent training. We designed a two-stage correlation strategy for plan optimization, including self-correlation and model preference correlation. Additionally, we utilized the generated plans to train the MultiMedia-Agent via a three stage approach including base/success plan finetune and preference optimization. The comparison results demonstrate that the our approaches are effective and the MultiMedia-Agent can generate better multimedia content compared to novel models.
title A Versatile Multimodal Agent for Multimedia Content Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.03250