Communicative Agents for Slideshow Storytelling Video Generation based on LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Jingxing, Shen, Jinrong, Yao, Yusheng, Wang, Shuangqing, Wang, Qian, Wang, Yuling
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908512295059456
author Fan, Jingxing
Shen, Jinrong
Yao, Yusheng
Wang, Shuangqing
Wang, Qian
Wang, Yuling
author_facet Fan, Jingxing
Shen, Jinrong
Yao, Yusheng
Wang, Shuangqing
Wang, Qian
Wang, Yuling
contents With the rapid advancement of artificial intelligence (AI), the proliferation of AI-generated content (AIGC) tasks has significantly accelerated developments in text-to-video generation. As a result, the field of video production is undergoing a transformative shift. However, conventional text-to-video models are typically constrained by high computational costs. In this study, we propose Video-Generation-Team (VGTeam), a novel slide show video generation system designed to redefine the video creation pipeline through the integration of large language models (LLMs). VGTeam is composed of a suite of communicative agents, each responsible for a distinct aspect of video generation, such as scriptwriting, scene creation, and audio design. These agents operate collaboratively within a chat tower workflow, transforming user-provided textual prompts into coherent, slide-style narrative videos. By emulating the sequential stages of traditional video production, VGTeam achieves remarkable improvements in both efficiency and scalability, while substantially reducing computational overhead. On average, the system generates videos at a cost of only $0.103, with a successful generation rate of 98.4%. Importantly, this framework maintains a high degree of creative fidelity and customization. The implications of VGTeam are far-reaching. It democratizes video production by enabling broader access to high-quality content creation without the need for extensive resources. Furthermore, it highlights the transformative potential of language models in creative domains and positions VGTeam as a pioneering system for next-generation content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
Fan, Jingxing
Shen, Jinrong
Yao, Yusheng
Wang, Shuangqing
Wang, Qian
Wang, Yuling
Artificial Intelligence
68T50, 68T42
I.2.6; I.2.7
With the rapid advancement of artificial intelligence (AI), the proliferation of AI-generated content (AIGC) tasks has significantly accelerated developments in text-to-video generation. As a result, the field of video production is undergoing a transformative shift. However, conventional text-to-video models are typically constrained by high computational costs. In this study, we propose Video-Generation-Team (VGTeam), a novel slide show video generation system designed to redefine the video creation pipeline through the integration of large language models (LLMs). VGTeam is composed of a suite of communicative agents, each responsible for a distinct aspect of video generation, such as scriptwriting, scene creation, and audio design. These agents operate collaboratively within a chat tower workflow, transforming user-provided textual prompts into coherent, slide-style narrative videos. By emulating the sequential stages of traditional video production, VGTeam achieves remarkable improvements in both efficiency and scalability, while substantially reducing computational overhead. On average, the system generates videos at a cost of only $0.103, with a successful generation rate of 98.4%. Importantly, this framework maintains a high degree of creative fidelity and customization. The implications of VGTeam are far-reaching. It democratizes video production by enabling broader access to high-quality content creation without the need for extensive resources. Furthermore, it highlights the transformative potential of language models in creative domains and positions VGTeam as a pioneering system for next-generation content creation.
title Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
topic Artificial Intelligence
68T50, 68T42
I.2.6; I.2.7
url https://arxiv.org/abs/2509.01277