Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Laakkonen, Janne, Kukanov, Ivan, Hautamäki, Ville
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912590544764928
author Laakkonen, Janne
Kukanov, Ivan
Hautamäki, Ville
author_facet Laakkonen, Janne
Kukanov, Ivan
Hautamäki, Ville
contents Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to novel deepfake methods not represented in training. To address this, we propose a mixture-of-LoRA-experts approach that integrates multiple low-rank adapters (LoRA) into the model's attention layers. A routing mechanism selectively activates specialized experts, enhancing adaptability to evolving deepfake attacks. Experimental results show that our method outperforms standard fine-tuning in both in-domain and out-of-domain scenarios, reducing equal error rates relative to baseline models. Notably, our best MoE-LoRA model lowers the average out-of-domain EER from 8.55\% to 6.08\%, demonstrating its effectiveness in achieving generalizable audio deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13878
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
Laakkonen, Janne
Kukanov, Ivan
Hautamäki, Ville
Audio and Speech Processing
Machine Learning
Sound
Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to novel deepfake methods not represented in training. To address this, we propose a mixture-of-LoRA-experts approach that integrates multiple low-rank adapters (LoRA) into the model's attention layers. A routing mechanism selectively activates specialized experts, enhancing adaptability to evolving deepfake attacks. Experimental results show that our method outperforms standard fine-tuning in both in-domain and out-of-domain scenarios, reducing equal error rates relative to baseline models. Notably, our best MoE-LoRA model lowers the average out-of-domain EER from 8.55\% to 6.08\%, demonstrating its effectiveness in achieving generalizable audio deepfake detection.
title Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2509.13878