Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chuchra, Akanksha, Reddy, Shukesh, Mishra, Sudeepta, Das, Abhijit, Dhall, Abhinav
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914230560620544
author Chuchra, Akanksha
Reddy, Shukesh
Mishra, Sudeepta
Das, Abhijit
Dhall, Abhinav
author_facet Chuchra, Akanksha
Reddy, Shukesh
Mishra, Sudeepta
Das, Abhijit
Dhall, Abhinav
contents While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we aim to explore the potential of MLLMs for audio deepfake detection. Combining audio inputs with a range of text prompts as queries to find out the viability of MLLMs to learn robust representations across modalities for audio deepfake detection. Therefore, we attempt to explore text-aware and context-rich, question-answer based prompts with binary decisions. We hypothesise that such a feature-guided reasoning will help in facilitating deeper multimodal understanding and enable robust feature learning for audio deepfake detection. We evaluate the performance of two MLLMs, Qwen2-Audio-7B-Instruct and SALMONN, in two evaluation modes: (a) zero-shot and (b) fine-tuned. Our experiments demonstrate that combining audio with a multi-prompt approach could be a viable way forward for audio deepfake detection. Our experiments show that the models perform poorly without task-specific training and struggle to generalise to out-of-domain data. However, they achieve good performance on in-domain data with minimal supervision, indicating promising potential for audio deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00777
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
Chuchra, Akanksha
Reddy, Shukesh
Mishra, Sudeepta
Das, Abhijit
Dhall, Abhinav
Sound
Computer Vision and Pattern Recognition
While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we aim to explore the potential of MLLMs for audio deepfake detection. Combining audio inputs with a range of text prompts as queries to find out the viability of MLLMs to learn robust representations across modalities for audio deepfake detection. Therefore, we attempt to explore text-aware and context-rich, question-answer based prompts with binary decisions. We hypothesise that such a feature-guided reasoning will help in facilitating deeper multimodal understanding and enable robust feature learning for audio deepfake detection. We evaluate the performance of two MLLMs, Qwen2-Audio-7B-Instruct and SALMONN, in two evaluation modes: (a) zero-shot and (b) fine-tuned. Our experiments demonstrate that combining audio with a multi-prompt approach could be a viable way forward for audio deepfake detection. Our experiments show that the models perform poorly without task-specific training and struggle to generalise to out-of-domain data. However, they achieve good performance on in-domain data with minimal supervision, indicating promising potential for audio deepfake detection.
title Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
topic Sound
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.00777