Towards generalisable and calibrated synthetic speech detection with self-supervised representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Pascu, Octavian, Stan, Adriana, Oneata, Dan, Oneata, Elisabeta, Cucu, Horia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
di: Pascu, Octavian, et al.
Pubblicazione: (2024)
di: Pascu, Octavian, et al.
Pubblicazione: (2024)
Echoes: A semantically-aligned music deepfake detection dataset
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
WavLM model ensemble for audio deepfake detection
di: Combei, David, et al.
Pubblicazione: (2024)
di: Combei, David, et al.
Pubblicazione: (2024)
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
di: Stan, Adriana, et al.
Pubblicazione: (2025)
di: Stan, Adriana, et al.
Pubblicazione: (2025)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
Unmasking real-world audio deepfakes: A data-centric approach
di: Combei, David, et al.
Pubblicazione: (2025)
di: Combei, David, et al.
Pubblicazione: (2025)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
Investigating self-supervised representations for audio-visual deepfake detection
di: Boldisor, Dragos-Alexandru, et al.
Pubblicazione: (2025)
di: Boldisor, Dragos-Alexandru, et al.
Pubblicazione: (2025)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
The mutual exclusivity bias of bilingual visually grounded speech models
di: Oneata, Dan, et al.
Pubblicazione: (2025)
di: Oneata, Dan, et al.
Pubblicazione: (2025)
AxLSTMs: learning self-supervised audio representations with xLSTMs
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
di: Lyth, Dan, et al.
Pubblicazione: (2024)
di: Lyth, Dan, et al.
Pubblicazione: (2024)
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
di: Wang, Bo, et al.
Pubblicazione: (2024)
di: Wang, Bo, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
A low latency attention module for streaming self-supervised speech representation learning
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
di: Striletchi, Vlad, et al.
Pubblicazione: (2024)
di: Striletchi, Vlad, et al.
Pubblicazione: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
di: Wang, Yi, et al.
Pubblicazione: (2025)
di: Wang, Yi, et al.
Pubblicazione: (2025)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
Self-supervised learning of speech representations with Dutch archival data
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts
di: Roquefort, Filomene, et al.
Pubblicazione: (2025)
di: Roquefort, Filomene, et al.
Pubblicazione: (2025)
Semantic enrichment towards efficient speech representations
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
Selfsupervised learning for pathological speech detection
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
On the relationship between speech and hearing
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
di: Pascu, Octavian, et al.
Pubblicazione: (2024) -
Echoes: A semantically-aligned music deepfake detection dataset
di: Pascu, Octavian, et al.
Pubblicazione: (2026) -
WavLM model ensemble for audio deepfake detection
di: Combei, David, et al.
Pubblicazione: (2024) -
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
di: Stan, Adriana, et al.
Pubblicazione: (2025) -
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)