Unmasking real-world audio deepfakes: A data-centric approach
Fuente:
arXiv
Saved in:
| Main Authors: | Combei, David, Stan, Adriana, Oneata, Dan, Müller, Nicolas, Cucu, Horia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WavLM model ensemble for audio deepfake detection
by: Combei, David, et al.
Published: (2024)
by: Combei, David, et al.
Published: (2024)
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
by: Stan, Adriana, et al.
Published: (2025)
by: Stan, Adriana, et al.
Published: (2025)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
Echoes: A semantically-aligned music deepfake detection dataset
by: Pascu, Octavian, et al.
Published: (2026)
by: Pascu, Octavian, et al.
Published: (2026)
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
by: Pascu, Octavian, et al.
Published: (2024)
by: Pascu, Octavian, et al.
Published: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
by: Pascu, Octavian, et al.
Published: (2023)
by: Pascu, Octavian, et al.
Published: (2023)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
by: Răgman, Teodora, et al.
Published: (2026)
by: Răgman, Teodora, et al.
Published: (2026)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Human perception of audio deepfakes: the role of language and speaking style
by: Segundo, Eugenia San, et al.
Published: (2025)
by: Segundo, Eugenia San, et al.
Published: (2025)
Open Source State-Of-the-Art Solution for Romanian Speech Recognition
by: Pirlogeanu, Gabriel, et al.
Published: (2025)
by: Pirlogeanu, Gabriel, et al.
Published: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
by: Combei, David
Published: (2025)
by: Combei, David
Published: (2025)
A robust audio deepfake detection system via multi-view feature
by: Yang, Yujie, et al.
Published: (2024)
by: Yang, Yujie, et al.
Published: (2024)
Translating speech with just images
by: Oneata, Dan, et al.
Published: (2024)
by: Oneata, Dan, et al.
Published: (2024)
Visually grounded few-shot word learning in low-resource settings
by: Nortje, Leanne, et al.
Published: (2023)
by: Nortje, Leanne, et al.
Published: (2023)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Forensic deepfake audio detection using segmental speech features
by: Yang, Tianle, et al.
Published: (2025)
by: Yang, Tianle, et al.
Published: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
by: Ratnarajah, Anton Jeran
Published: (2024)
by: Ratnarajah, Anton Jeran
Published: (2024)
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
by: Răgman, Teodora, et al.
Published: (2024)
by: Răgman, Teodora, et al.
Published: (2024)
The mutual exclusivity bias of bilingual visually grounded speech models
by: Oneata, Dan, et al.
Published: (2025)
by: Oneata, Dan, et al.
Published: (2025)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
by: Nortje, Leanne, et al.
Published: (2024)
by: Nortje, Leanne, et al.
Published: (2024)
Explaining deep learning models for spoofing and deepfake detection with SHapley Additive exPlanations
by: Ge, Wanying, et al.
Published: (2021)
by: Ge, Wanying, et al.
Published: (2021)
Online incremental learning for audio classification using a pretrained audio model
by: Mulimani, Manjunath, et al.
Published: (2025)
by: Mulimani, Manjunath, et al.
Published: (2025)
GRAM: Spatial general-purpose audio representation models for real-world applications
by: Yuksel, Goksenin, et al.
Published: (2025)
by: Yuksel, Goksenin, et al.
Published: (2025)
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
by: Ali, Hashim, et al.
Published: (2025)
by: Ali, Hashim, et al.
Published: (2025)
Positive and negative sampling strategies for self-supervised learning on audio-video data
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
by: Striletchi, Vlad, et al.
Published: (2024)
by: Striletchi, Vlad, et al.
Published: (2024)
Scaling up masked audio encoder learning for general audio classification
by: Dinkel, Heinrich, et al.
Published: (2024)
by: Dinkel, Heinrich, et al.
Published: (2024)
Cryfish: On deep audio analysis with Large Language Models
by: Mitrofanov, Anton, et al.
Published: (2025)
by: Mitrofanov, Anton, et al.
Published: (2025)
Multiple Hankel matrix rank minimization for audio inpainting
by: Záviška, Pavel, et al.
Published: (2023)
by: Záviška, Pavel, et al.
Published: (2023)
SCORE: Scaling audio generation using Standardized COmposite REwards
by: Jung, Jaemin, et al.
Published: (2025)
by: Jung, Jaemin, et al.
Published: (2025)
Detecting music deepfakes is easy but actually hard
by: Afchar, Darius, et al.
Published: (2024)
by: Afchar, Darius, et al.
Published: (2024)
LDCodec: A high quality neural audio codec with low-complexity decoder
by: Jiang, Jiawei, et al.
Published: (2025)
by: Jiang, Jiawei, et al.
Published: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
by: Han, Runduo, et al.
Published: (2024)
by: Han, Runduo, et al.
Published: (2024)
Blind estimation of audio effects using an auto-encoder approach and differentiable digital signal processing
by: Peladeau, Côme, et al.
Published: (2023)
by: Peladeau, Côme, et al.
Published: (2023)
PitchFlower: A flow-based neural audio codec with pitch controllability
by: Torres, Diego, et al.
Published: (2025)
by: Torres, Diego, et al.
Published: (2025)
Towards predicting binaural audio quality in listeners with normal and impaired hearing
by: Biberger, Thomas, et al.
Published: (2025)
by: Biberger, Thomas, et al.
Published: (2025)
Sound event detection with audio-text models and heterogeneous temporal annotations
by: Harju, Manu, et al.
Published: (2025)
by: Harju, Manu, et al.
Published: (2025)
AudioMorphix: Training-free audio editing with diffusion probabilistic models
by: Liang, Jinhua, et al.
Published: (2025)
by: Liang, Jinhua, et al.
Published: (2025)
Multi-label audio classification with a noisy zero-shot teacher
by: Braun, Sebastian, et al.
Published: (2024)
by: Braun, Sebastian, et al.
Published: (2024)
Towards audio language modeling -- an overview
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
Similar Items
-
WavLM model ensemble for audio deepfake detection
by: Combei, David, et al.
Published: (2024) -
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
by: Stan, Adriana, et al.
Published: (2025) -
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
by: Pîrlogeanu, Gabriel, et al.
Published: (2026) -
Echoes: A semantically-aligned music deepfake detection dataset
by: Pascu, Octavian, et al.
Published: (2026) -
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
by: Pascu, Octavian, et al.
Published: (2024)