Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
Fuente:
arXiv
Saved in:
| Main Authors: | S, Sridhar, A, Nithin, Rifath, Shakeel, Raj, Vasantha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
by: Chiu, Pin-Yen, et al.
Published: (2025)
by: Chiu, Pin-Yen, et al.
Published: (2025)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
by: Chen, Weiliang, et al.
Published: (2024)
by: Chen, Weiliang, et al.
Published: (2024)
Extreme Compression of Adaptive Neural Images
by: Hoshikawa, Leo, et al.
Published: (2024)
by: Hoshikawa, Leo, et al.
Published: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
Freehand Sketch Generation from Mechanical Components
by: Liao, Zhichao, et al.
Published: (2024)
by: Liao, Zhichao, et al.
Published: (2024)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
by: Wang, Xuanchen, et al.
Published: (2025)
by: Wang, Xuanchen, et al.
Published: (2025)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
by: He, Liu, et al.
Published: (2024)
by: He, Liu, et al.
Published: (2024)
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
by: Huang, Feizhen, et al.
Published: (2025)
by: Huang, Feizhen, et al.
Published: (2025)
HiSC4D: Human-centered interaction and 4D Scene Capture in Large-scale Space Using Wearable IMUs and LiDAR
by: Dai, Yudi, et al.
Published: (2024)
by: Dai, Yudi, et al.
Published: (2024)
Coral Model Generation from Single Images for Virtual Reality Applications
by: Fu, Jie, et al.
Published: (2024)
by: Fu, Jie, et al.
Published: (2024)
Instant3D: Instant Text-to-3D Generation
by: Li, Ming, et al.
Published: (2023)
by: Li, Ming, et al.
Published: (2023)
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
by: Hoe, Jiun Tian, et al.
Published: (2023)
by: Hoe, Jiun Tian, et al.
Published: (2023)
Neuro-Oracle: A Trajectory-Aware Agentic RAG Framework for Interpretable Epilepsy Surgical Prognosis
by: Aiersilan, Aizierjiang, et al.
Published: (2026)
by: Aiersilan, Aizierjiang, et al.
Published: (2026)
Lester: rotoscope animation through video object segmentation and tracking
by: Tous, Ruben
Published: (2024)
by: Tous, Ruben
Published: (2024)
Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
by: Sar, Ayan, et al.
Published: (2025)
by: Sar, Ayan, et al.
Published: (2025)
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
by: Lyu, Tianle, et al.
Published: (2025)
by: Lyu, Tianle, et al.
Published: (2025)
FlashSplat: 2D to 3D Gaussian Splatting Segmentation Solved Optimally
by: Shen, Qiuhong, et al.
Published: (2024)
by: Shen, Qiuhong, et al.
Published: (2024)
Seeing World Dynamics in a Nutshell
by: Shen, Qiuhong, et al.
Published: (2025)
by: Shen, Qiuhong, et al.
Published: (2025)
A Survey on 3D Gaussian Splatting
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
by: Singer, Assaf, et al.
Published: (2025)
by: Singer, Assaf, et al.
Published: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
by: Pham, Kien T., et al.
Published: (2025)
by: Pham, Kien T., et al.
Published: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
by: Xu, Chuanzhi, et al.
Published: (2026)
by: Xu, Chuanzhi, et al.
Published: (2026)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
by: Guo, Xin, et al.
Published: (2025)
by: Guo, Xin, et al.
Published: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
by: Han, Wei, et al.
Published: (2023)
by: Han, Wei, et al.
Published: (2023)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
by: Hu, Xiaowei, et al.
Published: (2024)
by: Hu, Xiaowei, et al.
Published: (2024)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
by: Lin, Jiantao, et al.
Published: (2025)
by: Lin, Jiantao, et al.
Published: (2025)
Cross-Scenario Deraining Adaptation with Unpaired Data: Superpixel Structural Priors and Multi-Stage Pseudo-Rain Synthesis
by: Zhao, Kangbo, et al.
Published: (2026)
by: Zhao, Kangbo, et al.
Published: (2026)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
by: Wu, Yuheng, et al.
Published: (2026)
by: Wu, Yuheng, et al.
Published: (2026)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
by: Wang, Yuze, et al.
Published: (2025)
by: Wang, Yuze, et al.
Published: (2025)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
by: Cha, SeungJu, et al.
Published: (2025)
by: Cha, SeungJu, et al.
Published: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
by: Chai, Zenghao, et al.
Published: (2024)
by: Chai, Zenghao, et al.
Published: (2024)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
by: Flynn, John, et al.
Published: (2026)
by: Flynn, John, et al.
Published: (2026)
Similar Items
-
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023) -
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
by: Chiu, Pin-Yen, et al.
Published: (2025) -
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024) -
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
by: Chen, Weiliang, et al.
Published: (2024) -
Extreme Compression of Adaptive Neural Images
by: Hoshikawa, Leo, et al.
Published: (2024)