A Comprehensive Survey on Generative AI for Video-to-Music Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Shulei, Wu, Songruoyao, Wang, Zihao, Li, Shuyu, Zhang, Kejun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
von: Ji, Shulei, et al.
Veröffentlicht: (2025)
von: Ji, Shulei, et al.
Veröffentlicht: (2025)
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
A Survey of Foundation Models for Music Understanding
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
From Sound to Sight: Towards AI-authored Music Videos
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
Cross-Modal Learning for Music-to-Music-Video Description Generation
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
Generative AI for Music and Audio
von: Dong, Hao-Wen
Veröffentlicht: (2024)
von: Dong, Hao-Wen
Veröffentlicht: (2024)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
von: Luo, Jing, et al.
Veröffentlicht: (2025)
von: Luo, Jing, et al.
Veröffentlicht: (2025)
BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
Vision-to-Music Generation: A Survey
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Speak the Art: A Direct Speech to Image Generation Framework
von: Saeed, Mariam, et al.
Veröffentlicht: (2025)
von: Saeed, Mariam, et al.
Veröffentlicht: (2025)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
von: Zhang, Kang, et al.
Veröffentlicht: (2025)
A Survey on Evaluation Metrics for Music Generation
von: Kader, Faria Binte, et al.
Veröffentlicht: (2025)
von: Kader, Faria Binte, et al.
Veröffentlicht: (2025)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
von: Li, Peike, et al.
Veröffentlicht: (2023)
von: Li, Peike, et al.
Veröffentlicht: (2023)
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
von: Ma, Menghe, et al.
Veröffentlicht: (2026)
von: Ma, Menghe, et al.
Veröffentlicht: (2026)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
CoComposer: LLM Multi-agent Collaborative Music Composition
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
Segment-Factorized Full-Song Generation on Symbolic Piano Music
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
von: Chen, Ping-Yi, et al.
Veröffentlicht: (2025)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
PodAgent: A Comprehensive Framework for Podcast Generation
von: Xiao, Yujia, et al.
Veröffentlicht: (2025)
von: Xiao, Yujia, et al.
Veröffentlicht: (2025)
Embedding Alignment in Code Generation for Audio
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
von: Ji, Shulei, et al.
Veröffentlicht: (2025) -
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
von: Ji, Shulei, et al.
Veröffentlicht: (2023) -
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
von: Zuo, Heda, et al.
Veröffentlicht: (2025) -
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
von: Wu, Junxian, et al.
Veröffentlicht: (2025) -
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
von: Li, Shuyu, et al.
Veröffentlicht: (2025)