HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mahapatra, Aurosweta, Ulgen, Ismail Rasim, Sisman, Berrak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Emotion Fool Anti-spoofing?
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2026)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2026)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
Universal Speech Content Factorization
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
von: Viakhirev, Ivan, et al.
Veröffentlicht: (2026)
von: Viakhirev, Ivan, et al.
Veröffentlicht: (2026)
Improving Short Utterance Anti-Spoofing with AASIST2
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2023)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2023)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Emotion Fool Anti-spoofing?
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025) -
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2026) -
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024) -
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025) -
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)