Understanding Audiovisual Deepfake Detection: Techniques, Challenges, Human Factors and Perceptual Insights
Fuente:
arXiv
Saved in:
| Main Authors: | Hashmi, Ammarah, Shahzad, Sahibzada Adil, Lin, Chia-Wen, Tsao, Yu, Wang, Hsin-Min |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
by: Hashmi, Ammarah, et al.
Published: (2024)
by: Hashmi, Ammarah, et al.
Published: (2024)
AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection
by: Hashmi, Ammarah, et al.
Published: (2023)
by: Hashmi, Ammarah, et al.
Published: (2023)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
by: Shahzad, Sahibzada Adil, et al.
Published: (2024)
by: Shahzad, Sahibzada Adil, et al.
Published: (2024)
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
by: Fu, Jingwen, et al.
Published: (2025)
by: Fu, Jingwen, et al.
Published: (2025)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)
by: Mingote, Victoria, et al.
Published: (2024)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024)
by: Yue, Xianghu, et al.
Published: (2024)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
by: Ragano, Alessandro, et al.
Published: (2024)
by: Ragano, Alessandro, et al.
Published: (2024)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
by: Wen, Liuyuan
Published: (2024)
by: Wen, Liuyuan
Published: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
by: Shimada, Kazuki, et al.
Published: (2025)
by: Shimada, Kazuki, et al.
Published: (2025)
Perceptual Video Quality Assessment: A Survey
by: Min, Xiongkuo, et al.
Published: (2024)
by: Min, Xiongkuo, et al.
Published: (2024)
Rip Current Detection in Nearshore Areas through UAV Video Analysis with Almost Local-Isometric Embedding Techniques on Sphere
by: Sun, Anchen, et al.
Published: (2023)
by: Sun, Anchen, et al.
Published: (2023)
Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective
by: Chen, Peilin, et al.
Published: (2024)
by: Chen, Peilin, et al.
Published: (2024)
Adaptive Wireless Image Semantic Transmission and Over-The-Air Testing
by: Ding, Jiarun, et al.
Published: (2024)
by: Ding, Jiarun, et al.
Published: (2024)
Learning Perceptual Representations for Gaming NR-VQA with Multi-Task FR Signals
by: Chen, Yu-Chih, et al.
Published: (2026)
by: Chen, Yu-Chih, et al.
Published: (2026)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
by: Sulun, Serkan, et al.
Published: (2025)
by: Sulun, Serkan, et al.
Published: (2025)
Perceptual Quality Optimization of Image Super-Resolution
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
ABC: Adaptive BayesNet Structure Learning for Computational Scalable Multi-task Image Compression
by: Zhang, Yufeng, et al.
Published: (2025)
by: Zhang, Yufeng, et al.
Published: (2025)
Symmetric Entropy-Constrained Video Coding for Machines
by: Sun, Yuxiao, et al.
Published: (2025)
by: Sun, Yuxiao, et al.
Published: (2025)
Perceptual Depth Quality Assessment of Stereoscopic Omnidirectional Images
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Fine-grained Image Quality Assessment for Perceptual Image Restoration
by: Sheng, Xiangfei, et al.
Published: (2025)
by: Sheng, Xiangfei, et al.
Published: (2025)
Transform and Entropy Coding in AV2
by: Nalci, Alican, et al.
Published: (2026)
by: Nalci, Alican, et al.
Published: (2026)
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
JND-Guided Light-Weight Neural Pre-Filter for Perceptual Image Coding
by: He, Chenlong, et al.
Published: (2025)
by: He, Chenlong, et al.
Published: (2025)
Perceptual Learned Image Compression via End-to-End JND-Based Optimization
by: Pakdaman, Farhad, et al.
Published: (2024)
by: Pakdaman, Farhad, et al.
Published: (2024)
Multimodal Confidence Modeling in Audio-Visual Quality Assessment
by: Mithila, Mayesha Maliha R., et al.
Published: (2026)
by: Mithila, Mayesha Maliha R., et al.
Published: (2026)
Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3
by: Artioli, Emanuele, et al.
Published: (2026)
by: Artioli, Emanuele, et al.
Published: (2026)
Robust Live Streaming over LEO Satellite Constellations: Measurement, Analysis, and Handover-Aware Adaptation
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Enhanced Template-based Intra Mode Derivation with Adaptive Block Vector Replacement
by: Zhang, Jiaqi, et al.
Published: (2025)
by: Zhang, Jiaqi, et al.
Published: (2025)
Tube-Structured Incremental Semantic HARQ for Generative Video Receivers
by: Wang, Xuesong, et al.
Published: (2026)
by: Wang, Xuesong, et al.
Published: (2026)
Rate-Quality or Energy-Quality Pareto Fronts for Adaptive Video Streaming?
by: Katsenou, Angeliki, et al.
Published: (2024)
by: Katsenou, Angeliki, et al.
Published: (2024)
Smaller is Better: Generative Models Can Power Short Video Preloading
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Camel: Frame-Level Bandwidth Estimation for Low-Latency Live Streaming under Video Bitrate Undershooting
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
NiMark: A Non-intrusive Watermarking Framework against Screen-shooting Attacks
by: Wu, Yufeng, et al.
Published: (2026)
by: Wu, Yufeng, et al.
Published: (2026)
Similar Items
-
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
by: Hashmi, Ammarah, et al.
Published: (2024) -
AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection
by: Hashmi, Ammarah, et al.
Published: (2023) -
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
by: Shahzad, Sahibzada Adil, et al.
Published: (2023) -
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026) -
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
by: Shahzad, Sahibzada Adil, et al.
Published: (2024)