WavLM model ensemble for audio deepfake detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Combei, David, Stan, Adriana, Oneata, Dan, Cucu, Horia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unmasking real-world audio deepfakes: A data-centric approach
di: Combei, David, et al.
Pubblicazione: (2025)
di: Combei, David, et al.
Pubblicazione: (2025)
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
di: Stan, Adriana, et al.
Pubblicazione: (2025)
di: Stan, Adriana, et al.
Pubblicazione: (2025)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
di: Pascu, Octavian, et al.
Pubblicazione: (2024)
di: Pascu, Octavian, et al.
Pubblicazione: (2024)
Echoes: A semantically-aligned music deepfake detection dataset
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
Adapting WavLM for Speech Emotion Recognition
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
di: Lin, Jingru, et al.
Pubblicazione: (2024)
di: Lin, Jingru, et al.
Pubblicazione: (2024)
Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier
di: Guo, Yinlin, et al.
Pubblicazione: (2023)
di: Guo, Yinlin, et al.
Pubblicazione: (2023)
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
di: Stourbe, Theophile, et al.
Pubblicazione: (2024)
di: Stourbe, Theophile, et al.
Pubblicazione: (2024)
XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
di: Zhang, Qishan, et al.
Pubblicazione: (2024)
di: Zhang, Qishan, et al.
Pubblicazione: (2024)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
di: Nakazawa, Kazushi
Pubblicazione: (2026)
di: Nakazawa, Kazushi
Pubblicazione: (2026)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
di: Zhang, Fan, et al.
Pubblicazione: (2023)
di: Zhang, Fan, et al.
Pubblicazione: (2023)
A robust audio deepfake detection system via multi-view feature
di: Yang, Yujie, et al.
Pubblicazione: (2024)
di: Yang, Yujie, et al.
Pubblicazione: (2024)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Human perception of audio deepfakes: the role of language and speaking style
di: Segundo, Eugenia San, et al.
Pubblicazione: (2025)
di: Segundo, Eugenia San, et al.
Pubblicazione: (2025)
Open Source State-Of-the-Art Solution for Romanian Speech Recognition
di: Pirlogeanu, Gabriel, et al.
Pubblicazione: (2025)
di: Pirlogeanu, Gabriel, et al.
Pubblicazione: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
di: Combei, David
Pubblicazione: (2025)
di: Combei, David
Pubblicazione: (2025)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
Forensic deepfake audio detection using segmental speech features
di: Yang, Tianle, et al.
Pubblicazione: (2025)
di: Yang, Tianle, et al.
Pubblicazione: (2025)
Explaining deep learning models for spoofing and deepfake detection with SHapley Additive exPlanations
di: Ge, Wanying, et al.
Pubblicazione: (2021)
di: Ge, Wanying, et al.
Pubblicazione: (2021)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
Are audio DeepFake detection models polyglots?
di: Marek, Bartłomiej, et al.
Pubblicazione: (2024)
di: Marek, Bartłomiej, et al.
Pubblicazione: (2024)
Sound event detection with audio-text models and heterogeneous temporal annotations
di: Harju, Manu, et al.
Pubblicazione: (2025)
di: Harju, Manu, et al.
Pubblicazione: (2025)
The mutual exclusivity bias of bilingual visually grounded speech models
di: Oneata, Dan, et al.
Pubblicazione: (2025)
di: Oneata, Dan, et al.
Pubblicazione: (2025)
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
Online incremental learning for audio classification using a pretrained audio model
di: Mulimani, Manjunath, et al.
Pubblicazione: (2025)
di: Mulimani, Manjunath, et al.
Pubblicazione: (2025)
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
di: Răgman, Teodora, et al.
Pubblicazione: (2024)
di: Răgman, Teodora, et al.
Pubblicazione: (2024)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
WavCraft: Audio Editing and Generation with Large Language Models
di: Liang, Jinhua, et al.
Pubblicazione: (2024)
di: Liang, Jinhua, et al.
Pubblicazione: (2024)
Generalizable speech deepfake detection via meta-learned LoRA
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
Towards audio language modeling -- an overview
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
di: Ali, Hashim, et al.
Pubblicazione: (2025)
di: Ali, Hashim, et al.
Pubblicazione: (2025)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
di: Kounadis-Bastian, Dionyssos, et al.
Pubblicazione: (2024)
di: Kounadis-Bastian, Dionyssos, et al.
Pubblicazione: (2024)
AudioMorphix: Training-free audio editing with diffusion probabilistic models
di: Liang, Jinhua, et al.
Pubblicazione: (2025)
di: Liang, Jinhua, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unmasking real-world audio deepfakes: A data-centric approach
di: Combei, David, et al.
Pubblicazione: (2025) -
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
di: Stan, Adriana, et al.
Pubblicazione: (2025) -
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026) -
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023) -
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
di: Pascu, Octavian, et al.
Pubblicazione: (2024)