DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Klemt, Marcel, Segna, Carlotta, Rohrbach, Anna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-Shot Fake Video Detection by Audio-Visual Consistency
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
di: Huang, Lian, et al.
Pubblicazione: (2024)
di: Huang, Lian, et al.
Pubblicazione: (2024)
Trusted Fake Audio Detection Based on Dirichlet Distribution
di: Ding, Chi, et al.
Pubblicazione: (2025)
di: Ding, Chi, et al.
Pubblicazione: (2025)
Analyzing the Impact of Splicing Artifacts in Partially Fake Speech Signals
di: Negroni, Viola, et al.
Pubblicazione: (2024)
di: Negroni, Viola, et al.
Pubblicazione: (2024)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
di: Li, Maomao, et al.
Pubblicazione: (2026)
di: Li, Maomao, et al.
Pubblicazione: (2026)
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
di: Zhang, Kang, et al.
Pubblicazione: (2025)
di: Zhang, Kang, et al.
Pubblicazione: (2025)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
di: Lee, Kyungbok, et al.
Pubblicazione: (2024)
di: Lee, Kyungbok, et al.
Pubblicazione: (2024)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
di: Li, Wenyu, et al.
Pubblicazione: (2025)
di: Li, Wenyu, et al.
Pubblicazione: (2025)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
di: Jiang, Yuxuan, et al.
Pubblicazione: (2025)
di: Jiang, Yuxuan, et al.
Pubblicazione: (2025)
Embedding Alignment in Code Generation for Audio
di: Kouteili, Sam, et al.
Pubblicazione: (2025)
di: Kouteili, Sam, et al.
Pubblicazione: (2025)
Retrieval-Augmented Text-to-Audio Generation
di: Yuan, Yi, et al.
Pubblicazione: (2023)
di: Yuan, Yi, et al.
Pubblicazione: (2023)
Neural Style Transfer for Audio Spectograms
di: Verma, Prateek, et al.
Pubblicazione: (2018)
di: Verma, Prateek, et al.
Pubblicazione: (2018)
Are audio DeepFake detection models polyglots?
di: Marek, Bartłomiej, et al.
Pubblicazione: (2024)
di: Marek, Bartłomiej, et al.
Pubblicazione: (2024)
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
di: Burchett-Vass, Rhys, et al.
Pubblicazione: (2024)
di: Burchett-Vass, Rhys, et al.
Pubblicazione: (2024)
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
di: Chen, Liangyu, et al.
Pubblicazione: (2024)
di: Chen, Liangyu, et al.
Pubblicazione: (2024)
Towards Generating Diverse Audio Captions via Adversarial Training
di: Mei, Xinhao, et al.
Pubblicazione: (2022)
di: Mei, Xinhao, et al.
Pubblicazione: (2022)
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
di: Bang, Hayeon, et al.
Pubblicazione: (2024)
di: Bang, Hayeon, et al.
Pubblicazione: (2024)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
di: Alonso-Jiménez, Pablo, et al.
Pubblicazione: (2024)
di: Alonso-Jiménez, Pablo, et al.
Pubblicazione: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
di: Yuan, Yi, et al.
Pubblicazione: (2023)
di: Yuan, Yi, et al.
Pubblicazione: (2023)
Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
di: Wen, Penghui, et al.
Pubblicazione: (2023)
di: Wen, Penghui, et al.
Pubblicazione: (2023)
SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
di: Yang, Zhe, et al.
Pubblicazione: (2024)
di: Yang, Zhe, et al.
Pubblicazione: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
di: Batlle-Roca, Roser, et al.
Pubblicazione: (2024)
di: Batlle-Roca, Roser, et al.
Pubblicazione: (2024)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
di: Wu, Junxian, et al.
Pubblicazione: (2025)
di: Wu, Junxian, et al.
Pubblicazione: (2025)
From Sound to Sight: Towards AI-authored Music Videos
di: Vitasovic, Leo, et al.
Pubblicazione: (2025)
di: Vitasovic, Leo, et al.
Pubblicazione: (2025)
PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
di: Hou, Yang, et al.
Pubblicazione: (2024)
di: Hou, Yang, et al.
Pubblicazione: (2024)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
di: Zuo, Heda, et al.
Pubblicazione: (2025)
di: Zuo, Heda, et al.
Pubblicazione: (2025)
IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection
di: Kumar, Abhay, et al.
Pubblicazione: (2025)
di: Kumar, Abhay, et al.
Pubblicazione: (2025)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
di: Liu, Haohe, et al.
Pubblicazione: (2023)
di: Liu, Haohe, et al.
Pubblicazione: (2023)
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
di: Zeng, Donghuo, et al.
Pubblicazione: (2024)
di: Zeng, Donghuo, et al.
Pubblicazione: (2024)
SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
di: Zhang, You, et al.
Pubblicazione: (2024)
di: Zhang, You, et al.
Pubblicazione: (2024)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
di: Huang, Zhiqi, et al.
Pubblicazione: (2024)
di: Huang, Zhiqi, et al.
Pubblicazione: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
di: Liu, Haohe, et al.
Pubblicazione: (2022)
di: Liu, Haohe, et al.
Pubblicazione: (2022)
SafeEar: Content Privacy-Preserving Audio Deepfake Detection
di: Li, Xinfeng, et al.
Pubblicazione: (2024)
di: Li, Xinfeng, et al.
Pubblicazione: (2024)
LoVA: Long-form Video-to-Audio Generation
di: Cheng, Xin, et al.
Pubblicazione: (2024)
di: Cheng, Xin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Zero-Shot Fake Video Detection by Audio-Visual Consistency
di: Li, Xiaolou, et al.
Pubblicazione: (2024) -
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
di: Huang, Lian, et al.
Pubblicazione: (2024) -
Trusted Fake Audio Detection Based on Dirichlet Distribution
di: Ding, Chi, et al.
Pubblicazione: (2025) -
Analyzing the Impact of Splicing Artifacts in Partially Fake Speech Signals
di: Negroni, Viola, et al.
Pubblicazione: (2024) -
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
di: Li, Maomao, et al.
Pubblicazione: (2026)