Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hou, Siyuan, Liu, Shansong, Yuan, Ruibin, Xue, Wei, Shan, Ying, Zhao, Mangsuo, Zhang, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Melody-Guided Music Generation
von: Wei, Shaopeng, et al.
Veröffentlicht: (2024)
von: Wei, Shaopeng, et al.
Veröffentlicht: (2024)
LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation
von: Baker, Tom, et al.
Veröffentlicht: (2025)
von: Baker, Tom, et al.
Veröffentlicht: (2025)
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation via Transformer VAE
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
REFFLY: Melody-Constrained Lyrics Editing Model
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
von: Meng, Hao, et al.
Veröffentlicht: (2026)
von: Meng, Hao, et al.
Veröffentlicht: (2026)
MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
von: Lu, Tongyu, et al.
Veröffentlicht: (2025)
von: Lu, Tongyu, et al.
Veröffentlicht: (2025)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
von: Wang, Ju-Chiang, et al.
Veröffentlicht: (2024)
Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
Small Tunes Transformer: Exploring Macro & Micro-Level Hierarchies for Skeleton-Conditioned Melody Generation
von: Lv, Yishan, et al.
Veröffentlicht: (2024)
von: Lv, Yishan, et al.
Veröffentlicht: (2024)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
Automatic Melody Reduction via Shortest Path Finding
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
Improving Musical Accompaniment Co-creation via Diffusion Transformers
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
von: Zhang, Yixiao
Veröffentlicht: (2024)
von: Zhang, Yixiao
Veröffentlicht: (2024)
Steer-by-prior Editing of Symbolic Music Loops
von: Jonason, Nicolas, et al.
Veröffentlicht: (2024)
von: Jonason, Nicolas, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
von: Sakpiboonchit, Siratish
Veröffentlicht: (2025)
von: Sakpiboonchit, Siratish
Veröffentlicht: (2025)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Melody-Guided Music Generation
von: Wei, Shaopeng, et al.
Veröffentlicht: (2024) -
LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation
von: Baker, Tom, et al.
Veröffentlicht: (2025) -
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
von: Wu, Shangda, et al.
Veröffentlicht: (2024) -
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025) -
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)