LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Junchen, Ge, Xuri, Zheng, Kaiwen, Karatzoglou, Alexandros, Arapakis, Ioannis, Xin, Xin, Ni, Yongxin, Jose, Joemon M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910000314580992
author Fu, Junchen
Ge, Xuri
Zheng, Kaiwen
Karatzoglou, Alexandros
Arapakis, Ioannis
Xin, Xin
Ni, Yongxin
Jose, Joemon M.
author_facet Fu, Junchen
Ge, Xuri
Zheng, Kaiwen
Karatzoglou, Alexandros
Arapakis, Ioannis
Xin, Xin
Ni, Yongxin
Jose, Joemon M.
contents In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs) to autonomously create viral micro-videos, a largely untapped potential that could shape the future of AI-driven content creation. To address this gap, this paper presents the first exploration of LLM-assisted popular micro-video generation (LLMPopcorn). We selected popcorn as the icon for this paper because it symbolizes leisure and entertainment, aligning with this study on leveraging LLMs as assistants for generating popular micro-videos that are often consumed during leisure time. Specifically, we empirically study the following research questions: (i) How can LLMs be effectively utilized to assist popular micro-video generation? (ii) To what extent can prompt-based enhancements optimize the LLM-generated content for higher popularity? (iii) How well do various LLMs and video generators perform in the popular micro-video generation task? Exploring these questions, we show that advanced LLMs like DeepSeek-V3 can generate micro-videos with popularity rivaling human content. Prompt enhancement further boosts results, while benchmarking highlights DeepSeek-V3 and R1 for LLMs, and LTX-Video and HunyuanVideo for video generation. This work advances AI-assisted micro-video creation and opens new research directions. The code is publicly available at https://github.com/GAIR-Lab/LLMPopcorn.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
Fu, Junchen
Ge, Xuri
Zheng, Kaiwen
Karatzoglou, Alexandros
Arapakis, Ioannis
Xin, Xin
Ni, Yongxin
Jose, Joemon M.
Computation and Language
Computer Vision and Pattern Recognition
In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs) to autonomously create viral micro-videos, a largely untapped potential that could shape the future of AI-driven content creation. To address this gap, this paper presents the first exploration of LLM-assisted popular micro-video generation (LLMPopcorn). We selected popcorn as the icon for this paper because it symbolizes leisure and entertainment, aligning with this study on leveraging LLMs as assistants for generating popular micro-videos that are often consumed during leisure time. Specifically, we empirically study the following research questions: (i) How can LLMs be effectively utilized to assist popular micro-video generation? (ii) To what extent can prompt-based enhancements optimize the LLM-generated content for higher popularity? (iii) How well do various LLMs and video generators perform in the popular micro-video generation task? Exploring these questions, we show that advanced LLMs like DeepSeek-V3 can generate micro-videos with popularity rivaling human content. Prompt enhancement further boosts results, while benchmarking highlights DeepSeek-V3 and R1 for LLMs, and LTX-Video and HunyuanVideo for video generation. This work advances AI-assisted micro-video creation and opens new research directions. The code is publicly available at https://github.com/GAIR-Lab/LLMPopcorn.
title LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.12945