ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chou, Benjamin, Zhu, Yi, Koppisetti, Surya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915943420002304
author Chou, Benjamin
Zhu, Yi
Koppisetti, Surya
author_facet Chou, Benjamin
Zhu, Yi
Koppisetti, Surya
contents Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel \textbf{I}n-\textbf{C}ontext \textbf{L}earning paradigm with comparison-guidance for \textbf{A}udio \textbf{D}eepfake detection (\textbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to $2\times$ relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
Chou, Benjamin
Zhu, Yi
Koppisetti, Surya
Sound
Computation and Language
Audio and Speech Processing
Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel \textbf{I}n-\textbf{C}ontext \textbf{L}earning paradigm with comparison-guidance for \textbf{A}udio \textbf{D}eepfake detection (\textbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to $2\times$ relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.
title ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2604.16749