Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dharmaratnakar, Abhishek, Ranganathan, Srivaths, Das, Debanshu, Sinha, Anushree
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914450262458368
author Dharmaratnakar, Abhishek
Ranganathan, Srivaths
Das, Debanshu
Sinha, Anushree
author_facet Dharmaratnakar, Abhishek
Ranganathan, Srivaths
Das, Debanshu
Sinha, Anushree
contents The domain of automatic video trailer generation is currently undergoing a profound paradigm shift, transitioning from heuristic-based extraction methods to deep generative synthesis. While early methodologies relied heavily on low-level feature engineering, visual saliency, and rule-based heuristics to select representative shots, recent advancements in Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), and diffusion-based video synthesis have enabled systems that not only identify key moments but also construct coherent, emotionally resonant narratives. This survey provides a comprehensive technical review of this evolution, with a specific focus on generative techniques including autoregressive Transformers, LLM-orchestrated pipelines, and text-to-video foundation models like OpenAI's Sora and Google's Veo. We analyze the architectural progression from Graph Convolutional Networks (GCNs) to Trailer Generation Transformers (TGT), evaluate the economic implications of automated content velocity on User-Generated Content (UGC) platforms, and discuss the ethical challenges posed by high-fidelity neural synthesis. By synthesizing insights from recent literature, this report establishes a new taxonomy for AI-driven trailer generation in the era of foundation models, suggesting that future promotional video systems will move beyond extractive selection toward controllable generative editing and semantic reconstruction of trailers.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04953
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity
Dharmaratnakar, Abhishek
Ranganathan, Srivaths
Das, Debanshu
Sinha, Anushree
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Information Retrieval
Multimedia
The domain of automatic video trailer generation is currently undergoing a profound paradigm shift, transitioning from heuristic-based extraction methods to deep generative synthesis. While early methodologies relied heavily on low-level feature engineering, visual saliency, and rule-based heuristics to select representative shots, recent advancements in Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), and diffusion-based video synthesis have enabled systems that not only identify key moments but also construct coherent, emotionally resonant narratives. This survey provides a comprehensive technical review of this evolution, with a specific focus on generative techniques including autoregressive Transformers, LLM-orchestrated pipelines, and text-to-video foundation models like OpenAI's Sora and Google's Veo. We analyze the architectural progression from Graph Convolutional Networks (GCNs) to Trailer Generation Transformers (TGT), evaluate the economic implications of automated content velocity on User-Generated Content (UGC) platforms, and discuss the ethical challenges posed by high-fidelity neural synthesis. By synthesizing insights from recent literature, this report establishes a new taxonomy for AI-driven trailer generation in the era of foundation models, suggesting that future promotional video systems will move beyond extractive selection toward controllable generative editing and semantic reconstruction of trailers.
title Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Information Retrieval
Multimedia
url https://arxiv.org/abs/2604.04953