DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wei, Sha, Binzhu, Luo, Dan, Yang, Jing, Wang, Zhuo, Fan, Fan, Wu, Zhiyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912527141568512
author Chen, Wei
Sha, Binzhu
Luo, Dan
Yang, Jing
Wang, Zhuo
Fan, Fan
Wu, Zhiyong
author_facet Chen, Wei
Sha, Binzhu
Luo, Dan
Yang, Jing
Wang, Zhuo
Fan, Fan
Wu, Zhiyong
contents Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
Chen, Wei
Sha, Binzhu
Luo, Dan
Yang, Jing
Wang, Zhuo
Fan, Fan
Wu, Zhiyong
Sound
Artificial Intelligence
Machine Learning
Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations.
title DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
topic Sound
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.05978