DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912527141568512 |
|---|---|
| author | Chen, Wei Sha, Binzhu Luo, Dan Yang, Jing Wang, Zhuo Fan, Fan Wu, Zhiyong |
| author_facet | Chen, Wei Sha, Binzhu Luo, Dan Yang, Jing Wang, Zhuo Fan, Fan Wu, Zhiyong |
| contents | Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_05978 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching Chen, Wei Sha, Binzhu Luo, Dan Yang, Jing Wang, Zhuo Fan, Fan Wu, Zhiyong Sound Artificial Intelligence Machine Learning Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations. |
| title | DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching |
| topic | Sound Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2508.05978 |