An End-to-End Approach for Chord-Conditioned Song Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Shuochen, Lei, Shun, Zhuo, Fan, Liu, Hangyu, Liu, Feng, Tang, Boshi, Huang, Qiaochu, Kang, Shiyin, Wu, Zhiyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SongCreator: Lyrics-based Universal Song Generation
von: Lei, Shun, et al.
Veröffentlicht: (2024)
von: Lei, Shun, et al.
Veröffentlicht: (2024)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation
von: Lin, Zhiwei, et al.
Veröffentlicht: (2024)
von: Lin, Zhiwei, et al.
Veröffentlicht: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
von: Tan, Wei, et al.
Veröffentlicht: (2025)
von: Tan, Wei, et al.
Veröffentlicht: (2025)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
von: Wu, Dapeng, et al.
Veröffentlicht: (2026)
Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
von: Munakata, Hokuto, et al.
Veröffentlicht: (2024)
von: Munakata, Hokuto, et al.
Veröffentlicht: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
von: Liu, Shuiyun, et al.
Veröffentlicht: (2024)
von: Liu, Shuiyun, et al.
Veröffentlicht: (2024)
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
von: Zhang, Huaicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Huaicheng, et al.
Veröffentlicht: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
von: Landini, Federico, et al.
Veröffentlicht: (2023)
von: Landini, Federico, et al.
Veröffentlicht: (2023)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
von: Lei, Shun, et al.
Veröffentlicht: (2025)
von: Lei, Shun, et al.
Veröffentlicht: (2025)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SongCreator: Lyrics-based Universal Song Generation
von: Lei, Shun, et al.
Veröffentlicht: (2024) -
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024) -
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
von: Dai, Dongyang, et al.
Veröffentlicht: (2025) -
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024) -
Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation
von: Lin, Zhiwei, et al.
Veröffentlicht: (2024)