Improving Short Utterance Anti-Spoofing with AASIST2

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuxiang, Lu, Jingze, Shang, Zengqiang, Wang, Wenchao, Zhang, Pengyuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913185097842688
author Zhang, Yuxiang
Lu, Jingze
Shang, Zengqiang
Wang, Wenchao
Zhang, Pengyuan
author_facet Zhang, Yuxiang
Lu, Jingze
Shang, Zengqiang
Wang, Wenchao
Zhang, Pengyuan
contents The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08279
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving Short Utterance Anti-Spoofing with AASIST2
Zhang, Yuxiang
Lu, Jingze
Shang, Zengqiang
Wang, Wenchao
Zhang, Pengyuan
Audio and Speech Processing
Sound
The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets.
title Improving Short Utterance Anti-Spoofing with AASIST2
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.08279