MFAAN: Unveiling Audio Deepfakes with a Multi-Feature Authenticity Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Krishnan, Karthik Sivarama, Krishnan, Koushik Sivarama
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914692760338432
author Krishnan, Karthik Sivarama
Krishnan, Koushik Sivarama
author_facet Krishnan, Karthik Sivarama
Krishnan, Koushik Sivarama
contents In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in misinformation campaigns. To address this threat, we introduce the Multi-Feature Audio Authenticity Network (MFAAN), an advanced architecture tailored for the detection of fabricated audio content. MFAAN incorporates multiple parallel paths designed to harness the strengths of different audio representations, including Mel-frequency cepstral coefficients (MFCC), linear-frequency cepstral coefficients (LFCC), and Chroma Short Time Fourier Transform (Chroma-STFT). By synergistically fusing these features, MFAAN achieves a nuanced understanding of audio content, facilitating robust differentiation between genuine and manipulated recordings. Preliminary evaluations of MFAAN on two benchmark datasets, 'In-the-Wild' Audio Deepfake Data and The Fake-or-Real Dataset, demonstrate its superior performance, achieving accuracies of 98.93% and 94.47% respectively. Such results not only underscore the efficacy of MFAAN but also highlight its potential as a pivotal tool in the ongoing battle against deepfake audio content.
format Preprint
id arxiv_https___arxiv_org_abs_2311_03509
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MFAAN: Unveiling Audio Deepfakes with a Multi-Feature Authenticity Network
Krishnan, Karthik Sivarama
Krishnan, Koushik Sivarama
Sound
Artificial Intelligence
Audio and Speech Processing
In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in misinformation campaigns. To address this threat, we introduce the Multi-Feature Audio Authenticity Network (MFAAN), an advanced architecture tailored for the detection of fabricated audio content. MFAAN incorporates multiple parallel paths designed to harness the strengths of different audio representations, including Mel-frequency cepstral coefficients (MFCC), linear-frequency cepstral coefficients (LFCC), and Chroma Short Time Fourier Transform (Chroma-STFT). By synergistically fusing these features, MFAAN achieves a nuanced understanding of audio content, facilitating robust differentiation between genuine and manipulated recordings. Preliminary evaluations of MFAAN on two benchmark datasets, 'In-the-Wild' Audio Deepfake Data and The Fake-or-Real Dataset, demonstrate its superior performance, achieving accuracies of 98.93% and 94.47% respectively. Such results not only underscore the efficacy of MFAAN but also highlight its potential as a pivotal tool in the ongoing battle against deepfake audio content.
title MFAAN: Unveiling Audio Deepfakes with a Multi-Feature Authenticity Network
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2311.03509