Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dvirniak, Artem, Kushnir, Evgeny, Tarasov, Dmitrii, Iudin, Artem, Kiriukhin, Oleg, Pautov, Mikhail, Korzh, Dmitrii, Rogov, Oleg Y.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911508269629440
author Dvirniak, Artem
Kushnir, Evgeny
Tarasov, Dmitrii
Iudin, Artem
Kiriukhin, Oleg
Pautov, Mikhail
Korzh, Dmitrii
Rogov, Oleg Y.
author_facet Dvirniak, Artem
Kushnir, Evgeny
Tarasov, Dmitrii
Iudin, Artem
Kiriukhin, Oleg
Pautov, Mikhail
Korzh, Dmitrii
Rogov, Oleg Y.
contents The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10725
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
Dvirniak, Artem
Kushnir, Evgeny
Tarasov, Dmitrii
Iudin, Artem
Kiriukhin, Oleg
Pautov, Mikhail
Korzh, Dmitrii
Rogov, Oleg Y.
Sound
Artificial Intelligence
The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
title Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2603.10725