X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Youngseo, Yun, Kwan, Hong, Seokhyeon, Cha, Sihun, Koo, Colette Suhjung, Noh, Junyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior
by: Kim, Gihoon, et al.
Published: (2024)
by: Kim, Gihoon, et al.
Published: (2024)
SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
by: Hong, Seokhyeon, et al.
Published: (2025)
by: Hong, Seokhyeon, et al.
Published: (2025)
Neural Face Skinning for Mesh-agnostic Facial Expression Cloning
by: Cha, Sihun, et al.
Published: (2025)
by: Cha, Sihun, et al.
Published: (2025)
AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models
by: Yun, Kwan, et al.
Published: (2025)
by: Yun, Kwan, et al.
Published: (2025)
LeGO: Leveraging a Surface Deformation Network for Animatable Stylized Face Generation with One Example
by: Yoon, Soyeon, et al.
Published: (2024)
by: Yoon, Soyeon, et al.
Published: (2024)
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
by: Jeon, Junhyuk, et al.
Published: (2026)
by: Jeon, Junhyuk, et al.
Published: (2026)
Deep Learning Based Facial Retargeting Using Local Patches
by: Choi, Yeonsoo, et al.
Published: (2026)
by: Choi, Yeonsoo, et al.
Published: (2026)
Representative Feature Extraction During Diffusion Process for Sketch Extraction with One Example
by: Yun, Kwan, et al.
Published: (2024)
by: Yun, Kwan, et al.
Published: (2024)
StyleID: A Perception-Aware Dataset and Metric for Stylization-Agnostic Facial Identity Recognition
by: Yun, Kwan, et al.
Published: (2026)
by: Yun, Kwan, et al.
Published: (2026)
Skinned Motion Retargeting with Spatially Adaptive Interaction Guidance
by: Choi, Soojin, et al.
Published: (2026)
by: Choi, Soojin, et al.
Published: (2026)
ASMR: Adaptive Skeleton‐Mesh Rigging and Skinning via 2D Generative Prior
by: Seokhyeon Hong, et al.
Published: (2025)
by: Seokhyeon Hong, et al.
Published: (2025)
ASMR: Adaptive Skeleton-Mesh Rigging and Skinning via 2D Generative Prior
by: Hong, Seokhyeon, et al.
Published: (2025)
by: Hong, Seokhyeon, et al.
Published: (2025)
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
by: Lee, Seungmi, et al.
Published: (2025)
by: Lee, Seungmi, et al.
Published: (2025)
LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization
by: Zeng, Bokang, et al.
Published: (2026)
by: Zeng, Bokang, et al.
Published: (2026)
Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection
by: Peng, Jielun, et al.
Published: (2026)
by: Peng, Jielun, et al.
Published: (2026)
FFaceNeRF: Few-shot Face Editing in Neural Radiance Fields
by: Yun, Kwan, et al.
Published: (2025)
by: Yun, Kwan, et al.
Published: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
by: Wei, Fangda, et al.
Published: (2026)
by: Wei, Fangda, et al.
Published: (2026)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
by: Klein, Nicholas, et al.
Published: (2025)
by: Klein, Nicholas, et al.
Published: (2025)
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Semi-Supervised Domain Adaptation for Wildfire Detection
by: Jang, JooYoung, et al.
Published: (2024)
by: Jang, JooYoung, et al.
Published: (2024)
Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
FLAMe: Federated Learning with Attention Mechanism using Spatio-Temporal Keypoint Transformers for Pedestrian Fall Detection in Smart Cities
by: Kim, Byeonghun, et al.
Published: (2024)
by: Kim, Byeonghun, et al.
Published: (2024)
AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
by: Kukanov, Ivan, et al.
Published: (2025)
by: Kukanov, Ivan, et al.
Published: (2025)
Semi-supervised reference-based sketch extraction using a contrastive learning framework
by: Seo, Chang Wook, et al.
Published: (2024)
by: Seo, Chang Wook, et al.
Published: (2024)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
Deepfake Detection with Spatio-Temporal Consistency and Attention
by: Chen, Yunzhuo, et al.
Published: (2025)
by: Chen, Yunzhuo, et al.
Published: (2025)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
by: Lin, Chun-Jung, et al.
Published: (2024)
by: Lin, Chun-Jung, et al.
Published: (2024)
Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition
by: Rajasekhar, G, et al.
Published: (2024)
by: Rajasekhar, G, et al.
Published: (2024)
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
by: Kim, Youngseo, et al.
Published: (2025)
by: Kim, Youngseo, et al.
Published: (2025)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
by: Oorloff, Trevine, et al.
Published: (2024)
by: Oorloff, Trevine, et al.
Published: (2024)
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
by: Astrid, Marcella, et al.
Published: (2025)
by: Astrid, Marcella, et al.
Published: (2025)
Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
by: Noh, Jeonghyun, et al.
Published: (2025)
by: Noh, Jeonghyun, et al.
Published: (2025)
Robust Deepfake Detection, NTIRE 2026 Challenge: Report
by: Hopf, Benedikt, et al.
Published: (2026)
by: Hopf, Benedikt, et al.
Published: (2026)
Practical Manipulation Model for Robust Deepfake Detection
by: Hopf, Benedikt, et al.
Published: (2025)
by: Hopf, Benedikt, et al.
Published: (2025)
Ensemble-Based Deepfake Detection using State-of-the-Art Models with Robust Cross-Dataset Generalisation
by: Wahab, Haroon, et al.
Published: (2025)
by: Wahab, Haroon, et al.
Published: (2025)
Similar Items
-
NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior
by: Kim, Gihoon, et al.
Published: (2024) -
SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
by: Hong, Seokhyeon, et al.
Published: (2025) -
Neural Face Skinning for Mesh-agnostic Facial Expression Cloning
by: Cha, Sihun, et al.
Published: (2025) -
AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models
by: Yun, Kwan, et al.
Published: (2025) -
LeGO: Leveraging a Surface Deformation Network for Animatable Stylized Face Generation with One Example
by: Yoon, Soyeon, et al.
Published: (2024)