How Does Instrumental Music Help SingFake Detection?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xuanjun, Hu, Chia-Yu, Lin, I-Ming, Lin, Yi-Cheng, Chiu, I-Hsiang, Zhang, You, Huang, Sung-Feng, Yang, Yi-Hsuan, Wu, Haibin, Lee, Hung-yi, Jang, Jyh-Shing Roger
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911385722552320
author Chen, Xuanjun
Hu, Chia-Yu
Lin, I-Ming
Lin, Yi-Cheng
Chiu, I-Hsiang
Zhang, You
Huang, Sung-Feng
Yang, Yi-Hsuan
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
author_facet Chen, Xuanjun
Hu, Chia-Yu
Lin, I-Ming
Lin, Yi-Cheng
Chiu, I-Hsiang
Zhang, You
Huang, Sung-Feng
Yang, Yi-Hsuan
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
contents Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different backbones, unpaired instrumental tracks, and frequency subbands. To analyze the representational effect, we probe how fine-tuning alters encoders' speech and music capabilities. Our results show that instrumental accompaniment acts mainly as data augmentation rather than providing intrinsic cues (e.g., rhythm or harmony). Furthermore, fine-tuning increases reliance on shallow speaker features while reducing sensitivity to content, paralinguistic, and semantic information. These insights clarify how models exploit vocal versus instrumental cues and can inform the design of more interpretable and robust SingFake detection systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Does Instrumental Music Help SingFake Detection?
Chen, Xuanjun
Hu, Chia-Yu
Lin, I-Ming
Lin, Yi-Cheng
Chiu, I-Hsiang
Zhang, You
Huang, Sung-Feng
Yang, Yi-Hsuan
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
Sound
Audio and Speech Processing
Signal Processing
Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different backbones, unpaired instrumental tracks, and frequency subbands. To analyze the representational effect, we probe how fine-tuning alters encoders' speech and music capabilities. Our results show that instrumental accompaniment acts mainly as data augmentation rather than providing intrinsic cues (e.g., rhythm or harmony). Furthermore, fine-tuning increases reliance on shallow speaker features while reducing sensitivity to content, paralinguistic, and semantic information. These insights clarify how models exploit vocal versus instrumental cues and can inform the design of more interpretable and robust SingFake detection systems.
title How Does Instrumental Music Help SingFake Detection?
topic Sound
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2509.14675