Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914006714810368 |
|---|---|
| author | N, Rishith Sadashiv T Bedge, Abhishek Bore, Saisha Suresh Mishra, Jagabandhu Bhattacharjee, Mrinmoy Prasanna, S R Mahadeva |
| author_facet | N, Rishith Sadashiv T Bedge, Abhishek Bore, Saisha Suresh Mishra, Jagabandhu Bhattacharjee, Mrinmoy Prasanna, S R Mahadeva |
| contents | Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_01034 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection N, Rishith Sadashiv T Bedge, Abhishek Bore, Saisha Suresh Mishra, Jagabandhu Bhattacharjee, Mrinmoy Prasanna, S R Mahadeva Audio and Speech Processing Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization. |
| title | Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.01034 |