Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Xin, Wanying, Ge, Yamagishi, Junichi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910039255547904
author Wang, Xin
Wanying, Ge
Yamagishi, Junichi
author_facet Wang, Xin
Wanying, Ge
Yamagishi, Junichi
contents Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuning paradigm using speech foundation models, most approaches rely solely on supervised fine-tuning (SFT). Inspired by the field of large language models, wherein reinforcement learning (RL) is used for model fine-tuning, we investigate the impact of RL, specifically Group Relative Policy Optimization (GRPO). The results from experiments using multiple detectors and test sets indicate that pure GRPO-based fine-tuning improves performance on out-of-domain test sets while maintaining performance on target-domain test data. This approach outperforms both SFT-only and hybrid setups. Our ablation studies further suggest that the negative reward in GRPO may be a key factor in this improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
Wang, Xin
Wanying, Ge
Yamagishi, Junichi
Audio and Speech Processing
Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuning paradigm using speech foundation models, most approaches rely solely on supervised fine-tuning (SFT). Inspired by the field of large language models, wherein reinforcement learning (RL) is used for model fine-tuning, we investigate the impact of RL, specifically Group Relative Policy Optimization (GRPO). The results from experiments using multiple detectors and test sets indicate that pure GRPO-based fine-tuning improves performance on out-of-domain test sets while maintaining performance on target-domain test data. This approach outperforms both SFT-only and hybrid setups. Our ablation studies further suggest that the negative reward in GRPO may be a key factor in this improvement.
title Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.02914