Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nallaguntla, Vamshi, Kshirsagar, Shruti, Avila, Anderson R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation
von: Nallaguntla, Vamshi, et al.
Veröffentlicht: (2026)
von: Nallaguntla, Vamshi, et al.
Veröffentlicht: (2026)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
von: Landau, Gilad, et al.
Veröffentlicht: (2025)
von: Landau, Gilad, et al.
Veröffentlicht: (2025)
Does Audio Deepfake Detection Generalize?
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Targeted Augmented Data for Audio Deepfake Detection
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
von: Zhang, Kuiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Kuiyuan, et al.
Veröffentlicht: (2024)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
von: Shan, Siyuan, et al.
Veröffentlicht: (2023)
von: Shan, Siyuan, et al.
Veröffentlicht: (2023)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Using Phonemes in cascaded S2S translation pipeline
von: Pilz, Rene, et al.
Veröffentlicht: (2025)
von: Pilz, Rene, et al.
Veröffentlicht: (2025)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
von: Channing, Georgia, et al.
Veröffentlicht: (2024)
von: Channing, Georgia, et al.
Veröffentlicht: (2024)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
Toward Fully Self-Supervised Multi-Pitch Estimation
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2024)
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2024)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
von: Miara, Victor, et al.
Veröffentlicht: (2024)
von: Miara, Victor, et al.
Veröffentlicht: (2024)
Music Emotion Prediction Using Recurrent Neural Networks
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
Self-Supervised Learning for Speaker Recognition: A study and review
von: Lepage, Theo, et al.
Veröffentlicht: (2026)
von: Lepage, Theo, et al.
Veröffentlicht: (2026)
Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
Singer Identity Representation Learning using Self-Supervised Techniques
von: Torres, Bernardo, et al.
Veröffentlicht: (2024)
von: Torres, Bernardo, et al.
Veröffentlicht: (2024)
Self-Supervised Learning for Few-Shot Bird Sound Classification
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
von: Ouyang, Qianhe
Veröffentlicht: (2025)
von: Ouyang, Qianhe
Veröffentlicht: (2025)
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
von: Lepage, Theo, et al.
Veröffentlicht: (2025)
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2025)
von: Cwitkowitz, Frank, et al.
Veröffentlicht: (2025)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
von: Kim, Seungmin, et al.
Veröffentlicht: (2025)
von: Kim, Seungmin, et al.
Veröffentlicht: (2025)
Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
von: Lepage, Théo, et al.
Veröffentlicht: (2022)
von: Lepage, Théo, et al.
Veröffentlicht: (2022)
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations
von: Heggan, Calum, et al.
Veröffentlicht: (2023)
von: Heggan, Calum, et al.
Veröffentlicht: (2023)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
Additive Margin in Contrastive Self-Supervised Frameworks to Learn Discriminative Speaker Representations
von: Lepage, Theo, et al.
Veröffentlicht: (2024)
von: Lepage, Theo, et al.
Veröffentlicht: (2024)
What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2025) -
Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024) -
PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation
von: Nallaguntla, Vamshi, et al.
Veröffentlicht: (2026) -
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
von: Landau, Gilad, et al.
Veröffentlicht: (2025) -
Does Audio Deepfake Detection Generalize?
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)