Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ou, Longshen, Ma, Xichu, Wang, Ye |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REFFLY: Melody-Constrained Lyrics Editing Model
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
von: Meng, Hao, et al.
Veröffentlicht: (2026)
von: Meng, Hao, et al.
Veröffentlicht: (2026)
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024)
Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization
von: Ou, Longshen, et al.
Veröffentlicht: (2024)
von: Ou, Longshen, et al.
Veröffentlicht: (2024)
Lead Instrument Detection from Multitrack Music
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
A Computational Analysis of Lyric Similarity Perception
von: Kim, Haven, et al.
Veröffentlicht: (2024)
von: Kim, Haven, et al.
Veröffentlicht: (2024)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Sing it, Narrate it: Quality Musical Lyrics Translation
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training
von: Yu, Jiaxing, et al.
Veröffentlicht: (2024)
von: Yu, Jiaxing, et al.
Veröffentlicht: (2024)
Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
von: Cai, Tracy, et al.
Veröffentlicht: (2024)
von: Cai, Tracy, et al.
Veröffentlicht: (2024)
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
von: Jafari, Farshad, et al.
Veröffentlicht: (2024)
von: Jafari, Farshad, et al.
Veröffentlicht: (2024)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
von: Chai, Li, et al.
Veröffentlicht: (2024)
von: Chai, Li, et al.
Veröffentlicht: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
von: Shi, Hao, et al.
Veröffentlicht: (2023)
von: Shi, Hao, et al.
Veröffentlicht: (2023)
Lyrics Transcription for Humans: A Readability-Aware Benchmark
von: Cífka, Ondřej, et al.
Veröffentlicht: (2024)
von: Cífka, Ondřej, et al.
Veröffentlicht: (2024)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
von: Ao, Junyi, et al.
Veröffentlicht: (2024)
von: Ao, Junyi, et al.
Veröffentlicht: (2024)
Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
von: Xu, Chun, et al.
Veröffentlicht: (2024)
von: Xu, Chun, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
Word-wise intonation model for cross-language TTS systems
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
von: Malan, Simon, et al.
Veröffentlicht: (2024)
von: Malan, Simon, et al.
Veröffentlicht: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
von: Malan, Simon, et al.
Veröffentlicht: (2025)
von: Malan, Simon, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
REFFLY: Melody-Constrained Lyrics Editing Model
von: Zhao, Songyan, et al.
Veröffentlicht: (2024) -
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
von: Meng, Hao, et al.
Veröffentlicht: (2026) -
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
von: Ding, Shuangrui, et al.
Veröffentlicht: (2024) -
Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization
von: Ou, Longshen, et al.
Veröffentlicht: (2024) -
Lead Instrument Detection from Multitrack Music
von: Ou, Longshen, et al.
Veröffentlicht: (2025)