REFFLY: Melody-Constrained Lyrics Editing Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Songyan, Li, Bingxuan, Tian, Yufei, Peng, Nanyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation
di: Ou, Longshen, et al.
Pubblicazione: (2023)
di: Ou, Longshen, et al.
Pubblicazione: (2023)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
di: Meng, Hao, et al.
Pubblicazione: (2026)
di: Meng, Hao, et al.
Pubblicazione: (2026)
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
di: Ding, Shuangrui, et al.
Pubblicazione: (2024)
di: Ding, Shuangrui, et al.
Pubblicazione: (2024)
Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
di: Zhuo, Le, et al.
Pubblicazione: (2023)
di: Zhuo, Le, et al.
Pubblicazione: (2023)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
di: Huang, Jiawen, et al.
Pubblicazione: (2024)
di: Huang, Jiawen, et al.
Pubblicazione: (2024)
A Computational Analysis of Lyric Similarity Perception
di: Kim, Haven, et al.
Pubblicazione: (2024)
di: Kim, Haven, et al.
Pubblicazione: (2024)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
di: Hao, Chunbo, et al.
Pubblicazione: (2026)
di: Hao, Chunbo, et al.
Pubblicazione: (2026)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
di: Hou, Siyuan, et al.
Pubblicazione: (2024)
di: Hou, Siyuan, et al.
Pubblicazione: (2024)
Sing it, Narrate it: Quality Musical Lyrics Translation
di: Ye, Zhuorui, et al.
Pubblicazione: (2024)
di: Ye, Zhuorui, et al.
Pubblicazione: (2024)
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
di: Jafari, Farshad, et al.
Pubblicazione: (2024)
di: Jafari, Farshad, et al.
Pubblicazione: (2024)
Not that Groove: Zero-Shot Symbolic Music Editing
di: Zhang, Li
Pubblicazione: (2025)
di: Zhang, Li
Pubblicazione: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
di: Kulshreshtha, Devang, et al.
Pubblicazione: (2024)
di: Kulshreshtha, Devang, et al.
Pubblicazione: (2024)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
di: Chai, Li, et al.
Pubblicazione: (2024)
di: Chai, Li, et al.
Pubblicazione: (2024)
Lyrics Transcription for Humans: A Readability-Aware Benchmark
di: Cífka, Ondřej, et al.
Pubblicazione: (2024)
di: Cífka, Ondřej, et al.
Pubblicazione: (2024)
Speech Editing -- a Summary
di: Kässmann, Tobias, et al.
Pubblicazione: (2024)
di: Kässmann, Tobias, et al.
Pubblicazione: (2024)
Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
di: Cai, Tracy, et al.
Pubblicazione: (2024)
di: Cai, Tracy, et al.
Pubblicazione: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training
di: Yu, Jiaxing, et al.
Pubblicazione: (2024)
di: Yu, Jiaxing, et al.
Pubblicazione: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
di: Wang, Peng, et al.
Pubblicazione: (2023)
di: Wang, Peng, et al.
Pubblicazione: (2023)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
di: Peng, Yifan, et al.
Pubblicazione: (2025)
di: Peng, Yifan, et al.
Pubblicazione: (2025)
Next Tokens Denoising for Speech Synthesis
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
di: Zhao, Zhixian, et al.
Pubblicazione: (2025)
di: Zhao, Zhixian, et al.
Pubblicazione: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
di: Li, Yuang, et al.
Pubblicazione: (2024)
di: Li, Yuang, et al.
Pubblicazione: (2024)
OpusLM: A Family of Open Unified Speech Language Models
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
di: Geng, Xuelong, et al.
Pubblicazione: (2025)
di: Geng, Xuelong, et al.
Pubblicazione: (2025)
Automatic Melody Reduction via Shortest Path Finding
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
di: Tang, Changli, et al.
Pubblicazione: (2023)
di: Tang, Changli, et al.
Pubblicazione: (2023)
Word-Level ASR Quality Estimation for Efficient Corpus Sampling and Post-Editing through Analyzing Attentions of a Reference-Free Metric
di: Javadi, Golara, et al.
Pubblicazione: (2024)
di: Javadi, Golara, et al.
Pubblicazione: (2024)
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
di: Li, Bohan, et al.
Pubblicazione: (2025)
di: Li, Bohan, et al.
Pubblicazione: (2025)
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
di: Ling, Shaoshi, et al.
Pubblicazione: (2023)
di: Ling, Shaoshi, et al.
Pubblicazione: (2023)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
di: Wu, Shangda, et al.
Pubblicazione: (2024)
di: Wu, Shangda, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation
di: Ou, Longshen, et al.
Pubblicazione: (2023) -
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
di: Meng, Hao, et al.
Pubblicazione: (2026) -
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
di: Ding, Shuangrui, et al.
Pubblicazione: (2024) -
Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
di: Li, Ruiqi, et al.
Pubblicazione: (2024) -
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
di: Zhuo, Le, et al.
Pubblicazione: (2023)