Towards the Next Frontier in Speech Representation Learning Using Disentanglement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krishna, Varun, Ganapathy, Sriram |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
Innovative Speech-Based Deep Learning Approaches for Parkinson's Disease Classification: A Systematic Review
von: van Gelderen, Lisanne, et al.
Veröffentlicht: (2024)
von: van Gelderen, Lisanne, et al.
Veröffentlicht: (2024)
Learning Disentangled Speech Representations
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
von: Zhu, Haina, et al.
Veröffentlicht: (2025)
von: Zhu, Haina, et al.
Veröffentlicht: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Real-time Speech Summarization for Medical Conversations
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
On the Role of Speech Data in Reducing Toxicity Detection Bias
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
Scaling Rich Style-Prompted Text-to-Speech Datasets
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
von: Bataev, Vladimir
Veröffentlicht: (2025)
von: Bataev, Vladimir
Veröffentlicht: (2025)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
von: Jia, Dongya, et al.
Veröffentlicht: (2025)
von: Jia, Dongya, et al.
Veröffentlicht: (2025)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization
von: Samin, Md. Nazmus Sadat, et al.
Veröffentlicht: (2024)
von: Samin, Md. Nazmus Sadat, et al.
Veröffentlicht: (2024)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
von: Li, Zehua Kcriss, et al.
Veröffentlicht: (2024)
von: Li, Zehua Kcriss, et al.
Veröffentlicht: (2024)
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
von: R, Aron, et al.
Veröffentlicht: (2024)
von: R, Aron, et al.
Veröffentlicht: (2024)
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
von: Xie, Tianxin, et al.
Veröffentlicht: (2024)
von: Xie, Tianxin, et al.
Veröffentlicht: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025) -
Improving Self-supervised Pre-training using Accent-Specific Codebooks
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024) -
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024) -
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
von: Lian, Jiachen, et al.
Veröffentlicht: (2022) -
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)