MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Boudaghi, Ali, Zare, Hadi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
por: Zhang, Yixiao, et al.
Publicado: (2024)
por: Zhang, Yixiao, et al.
Publicado: (2024)
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
por: Ji, Shulei, et al.
Publicado: (2023)
por: Ji, Shulei, et al.
Publicado: (2023)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
por: Zhang, Yixiao, et al.
Publicado: (2024)
por: Zhang, Yixiao, et al.
Publicado: (2024)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
por: Song, Jiahao, et al.
Publicado: (2025)
por: Song, Jiahao, et al.
Publicado: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
por: Niu, Xinlei, et al.
Publicado: (2025)
por: Niu, Xinlei, et al.
Publicado: (2025)
JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
por: Li, Peike, et al.
Publicado: (2023)
por: Li, Peike, et al.
Publicado: (2023)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
por: Zhao, Zijian, et al.
Publicado: (2025)
por: Zhao, Zijian, et al.
Publicado: (2025)
Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation
por: Zhu, Tingyu, et al.
Publicado: (2024)
por: Zhu, Tingyu, et al.
Publicado: (2024)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
por: Shao, Keren, et al.
Publicado: (2025)
por: Shao, Keren, et al.
Publicado: (2025)
BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
por: Luo, Jing, et al.
Publicado: (2024)
por: Luo, Jing, et al.
Publicado: (2024)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
por: Wang, Jinting, et al.
Publicado: (2025)
por: Wang, Jinting, et al.
Publicado: (2025)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
por: Retkowski, Jan, et al.
Publicado: (2024)
por: Retkowski, Jan, et al.
Publicado: (2024)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
por: Choi, Suhwan, et al.
Publicado: (2025)
por: Choi, Suhwan, et al.
Publicado: (2025)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
por: Batlle-Roca, Roser, et al.
Publicado: (2024)
por: Batlle-Roca, Roser, et al.
Publicado: (2024)
Generative AI for Music and Audio
por: Dong, Hao-Wen
Publicado: (2024)
por: Dong, Hao-Wen
Publicado: (2024)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
por: Long, Phillip, et al.
Publicado: (2024)
por: Long, Phillip, et al.
Publicado: (2024)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
por: Liu, Renhang, et al.
Publicado: (2024)
por: Liu, Renhang, et al.
Publicado: (2024)
A Survey of Foundation Models for Music Understanding
por: Li, Wenjun, et al.
Publicado: (2024)
por: Li, Wenjun, et al.
Publicado: (2024)
CoComposer: LLM Multi-agent Collaborative Music Composition
por: Xing, Peiwen, et al.
Publicado: (2025)
por: Xing, Peiwen, et al.
Publicado: (2025)
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
por: Wu, Junxian, et al.
Publicado: (2025)
por: Wu, Junxian, et al.
Publicado: (2025)
From Sound to Sight: Towards AI-authored Music Videos
por: Vitasovic, Leo, et al.
Publicado: (2025)
por: Vitasovic, Leo, et al.
Publicado: (2025)
LM2D: Lyrics- and Music-Driven Dance Synthesis
por: Yin, Wenjie, et al.
Publicado: (2024)
por: Yin, Wenjie, et al.
Publicado: (2024)
Segment-Factorized Full-Song Generation on Symbolic Piano Music
por: Chen, Ping-Yi, et al.
Publicado: (2025)
por: Chen, Ping-Yi, et al.
Publicado: (2025)
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
por: Alonso-Jiménez, Pablo, et al.
Publicado: (2024)
por: Alonso-Jiménez, Pablo, et al.
Publicado: (2024)
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
por: Ma, Menghe, et al.
Publicado: (2026)
por: Ma, Menghe, et al.
Publicado: (2026)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
por: Zuo, Heda, et al.
Publicado: (2025)
por: Zuo, Heda, et al.
Publicado: (2025)
The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
por: Nagarajan, Ashwin, et al.
Publicado: (2025)
por: Nagarajan, Ashwin, et al.
Publicado: (2025)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
por: Yuan, Ruibin, et al.
Publicado: (2025)
por: Yuan, Ruibin, et al.
Publicado: (2025)
Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
por: Luo, Jing, et al.
Publicado: (2025)
por: Luo, Jing, et al.
Publicado: (2025)
Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform
por: Telila, Yohannis, et al.
Publicado: (2025)
por: Telila, Yohannis, et al.
Publicado: (2025)
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
por: Ma, Yinghao, et al.
Publicado: (2026)
por: Ma, Yinghao, et al.
Publicado: (2026)
Music Enhancement with Deep Filters: A Technical Report for The ICASSP 2024 Cadenza Challenge
por: Shao, Keren, et al.
Publicado: (2024)
por: Shao, Keren, et al.
Publicado: (2024)
MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
por: Tan, Hao Hao, et al.
Publicado: (2024)
por: Tan, Hao Hao, et al.
Publicado: (2024)
Retrieval-Augmented Text-to-Audio Generation
por: Yuan, Yi, et al.
Publicado: (2023)
por: Yuan, Yi, et al.
Publicado: (2023)
Audio Transformers
por: Verma, Prateek, et al.
Publicado: (2021)
por: Verma, Prateek, et al.
Publicado: (2021)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
por: Hong, Fa-Ting, et al.
Publicado: (2024)
por: Hong, Fa-Ting, et al.
Publicado: (2024)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
por: Deng, Zihao, et al.
Publicado: (2023)
por: Deng, Zihao, et al.
Publicado: (2023)
Fast Text-to-Audio Generation with Adversarial Post-Training
por: Novack, Zachary, et al.
Publicado: (2025)
por: Novack, Zachary, et al.
Publicado: (2025)
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
por: Bang, Hayeon, et al.
Publicado: (2024)
por: Bang, Hayeon, et al.
Publicado: (2024)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
por: Lam, Max W. Y., et al.
Publicado: (2025)
por: Lam, Max W. Y., et al.
Publicado: (2025)
Ejemplares similares
-
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
por: Zhang, Yixiao, et al.
Publicado: (2024) -
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
por: Ji, Shulei, et al.
Publicado: (2023) -
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
por: Zhang, Yixiao, et al.
Publicado: (2024) -
MusFlow: Multimodal Music Generation via Conditional Flow Matching
por: Song, Jiahao, et al.
Publicado: (2025) -
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
por: Niu, Xinlei, et al.
Publicado: (2025)