Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Chuchra, Akanksha, Reddy, Shukesh, Mishra, Sudeepta, Das, Abhijit, Dhall, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
A Backbone Benchmarking Study on Self-supervised Learning as a Auxiliary Task with Texture-based Local Descriptors for Face Analysis
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
by: Yee, Phyo Thet, et al.
Published: (2025)
by: Yee, Phyo Thet, et al.
Published: (2025)
Exploring Green AI for Audio Deepfake Detection
by: Saha, Subhajit, et al.
Published: (2024)
by: Saha, Subhajit, et al.
Published: (2024)
Self-supervised Auxiliary Learning for Texture and Model-based Hybrid Robust and Fair Featuring in Face Analysis
by: Reddy, Shukesh, et al.
Published: (2024)
by: Reddy, Shukesh, et al.
Published: (2024)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
by: PN, Aravinda Reddy, et al.
Published: (2024)
by: PN, Aravinda Reddy, et al.
Published: (2024)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
by: Guo, Yuxin, et al.
Published: (2025)
by: Guo, Yuxin, et al.
Published: (2025)
Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion
by: Sutharya, S., et al.
Published: (2026)
by: Sutharya, S., et al.
Published: (2026)
TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models
by: Khan, Awais, et al.
Published: (2026)
by: Khan, Awais, et al.
Published: (2026)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
by: Oorloff, Trevine, et al.
Published: (2024)
by: Oorloff, Trevine, et al.
Published: (2024)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
by: Klein, Nicholas, et al.
Published: (2025)
by: Klein, Nicholas, et al.
Published: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
by: Tian, Zeyue, et al.
Published: (2026)
by: Tian, Zeyue, et al.
Published: (2026)
Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
by: Uddin, Kutub, et al.
Published: (2025)
by: Uddin, Kutub, et al.
Published: (2025)
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
by: Astrid, Marcella, et al.
Published: (2025)
by: Astrid, Marcella, et al.
Published: (2025)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
Statistics-aware Audio-visual Deepfake Detector
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
by: Zaman, Sayeem Been, et al.
Published: (2025)
by: Zaman, Sayeem Been, et al.
Published: (2025)
Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
by: Coletta, Emma, et al.
Published: (2025)
by: Coletta, Emma, et al.
Published: (2025)
Graph Propagated Projection Unlearning: A Unified Framework for Vision and Audio Discriminative Models
by: Pathak, Shreyansh, et al.
Published: (2026)
by: Pathak, Shreyansh, et al.
Published: (2026)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
by: Chen, Xuanjun, et al.
Published: (2025)
by: Chen, Xuanjun, et al.
Published: (2025)
End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Shared Multi-modal Embedding Space for Face-Voice Association
by: Simic, Christopher, et al.
Published: (2025)
by: Simic, Christopher, et al.
Published: (2025)
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
by: Li, Bingliang, et al.
Published: (2024)
by: Li, Bingliang, et al.
Published: (2024)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Aligned Better, Listen Better for Audio-Visual Large Language Models
by: Guo, Yuxin, et al.
Published: (2025)
by: Guo, Yuxin, et al.
Published: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
by: Dahan, Aviad, et al.
Published: (2026)
by: Dahan, Aviad, et al.
Published: (2026)
Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
by: Pianese, Alessandro, et al.
Published: (2024)
by: Pianese, Alessandro, et al.
Published: (2024)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
by: Yang, Jianxuan, et al.
Published: (2025)
by: Yang, Jianxuan, et al.
Published: (2025)
Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space
by: Chandra, Aashish, et al.
Published: (2026)
by: Chandra, Aashish, et al.
Published: (2026)
1M-Deepfakes Detection Challenge
by: Cai, Zhixi, et al.
Published: (2024)
by: Cai, Zhixi, et al.
Published: (2024)
Similar Items
-
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026) -
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025) -
A Backbone Benchmarking Study on Self-supervised Learning as a Auxiliary Task with Texture-based Local Descriptors for Face Analysis
by: Reddy, Shukesh, et al.
Published: (2026) -
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
by: Yee, Phyo Thet, et al.
Published: (2025) -
Exploring Green AI for Audio Deepfake Detection
by: Saha, Subhajit, et al.
Published: (2024)