Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: S, Sridhar, A, Nithin, Rifath, Shakeel, Raj, Vasantha
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913890008301568
author S, Sridhar
A, Nithin
Rifath, Shakeel
Raj, Vasantha
author_facet S, Sridhar
A, Nithin
Rifath, Shakeel
Raj, Vasantha
contents Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable Diffusion for high-fidelity image synthesis, GPT-2 for narrative structuring, and a hybrid audio pipeline using gTTS and YouTube-sourced music. It uses a five-scene framework, which is augmented by linear frame interpolation, cinematic post-processing (e.g., sharpening), and audio-video synchronization to provide professional-quality results. It was created in a GPU-accelerated Google Colab environment using Python 3.11. It has a dual-mode Gradio interface (Simple and Advanced), which supports resolutions of up to 1024x768 and frame rates of 15-30 FPS. Optimizations such as CUDA memory management and error handling ensure reliability. The experiments demonstrate outstanding visual quality, narrative coherence, and efficiency, furthering text-to-video synthesis for creative, educational, and industrial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
S, Sridhar
A, Nithin
Rifath, Shakeel
Raj, Vasantha
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Graphics
Multimedia
Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable Diffusion for high-fidelity image synthesis, GPT-2 for narrative structuring, and a hybrid audio pipeline using gTTS and YouTube-sourced music. It uses a five-scene framework, which is augmented by linear frame interpolation, cinematic post-processing (e.g., sharpening), and audio-video synchronization to provide professional-quality results. It was created in a GPU-accelerated Google Colab environment using Python 3.11. It has a dual-mode Gradio interface (Simple and Advanced), which supports resolutions of up to 1024x768 and frame rates of 15-30 FPS. Optimizations such as CUDA memory management and error handling ensure reliability. The experiments demonstrate outstanding visual quality, narrative coherence, and efficiency, furthering text-to-video synthesis for creative, educational, and industrial applications.
title Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Graphics
Multimedia
url https://arxiv.org/abs/2506.10005