Towards the Next Frontier in Speech Representation Learning Using Disentanglement
Fuente:
arXiv
Guardado en:
| Autores principales: | Krishna, Varun, Ganapathy, Sriram |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
por: Bhattacharya, Debarpan, et al.
Publicado: (2025)
por: Bhattacharya, Debarpan, et al.
Publicado: (2025)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
por: Prabhu, Darshan, et al.
Publicado: (2024)
por: Prabhu, Darshan, et al.
Publicado: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
por: Dutta, Soumya, et al.
Publicado: (2024)
por: Dutta, Soumya, et al.
Publicado: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
por: Lian, Jiachen, et al.
Publicado: (2022)
por: Lian, Jiachen, et al.
Publicado: (2022)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
por: Mu, Zhaoxi, et al.
Publicado: (2023)
por: Mu, Zhaoxi, et al.
Publicado: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
por: Seth, Ashish, et al.
Publicado: (2024)
por: Seth, Ashish, et al.
Publicado: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
por: Mohebbi, Hosein, et al.
Publicado: (2024)
por: Mohebbi, Hosein, et al.
Publicado: (2024)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
por: Elmakies, Avishai, et al.
Publicado: (2025)
por: Elmakies, Avishai, et al.
Publicado: (2025)
Towards Robust Speech Representation Learning for Thousands of Languages
por: Chen, William, et al.
Publicado: (2024)
por: Chen, William, et al.
Publicado: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
por: Rufai, Amina Mardiyyah, et al.
Publicado: (2020)
por: Rufai, Amina Mardiyyah, et al.
Publicado: (2020)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
por: Ye, Zhen, et al.
Publicado: (2024)
por: Ye, Zhen, et al.
Publicado: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
por: Elmakies, Avishai, et al.
Publicado: (2025)
por: Elmakies, Avishai, et al.
Publicado: (2025)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
por: Peng, Puyuan, et al.
Publicado: (2024)
por: Peng, Puyuan, et al.
Publicado: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
por: Singh, Prachi, et al.
Publicado: (2024)
por: Singh, Prachi, et al.
Publicado: (2024)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
por: Ju, Zeqian, et al.
Publicado: (2024)
por: Ju, Zeqian, et al.
Publicado: (2024)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Innovative Speech-Based Deep Learning Approaches for Parkinson's Disease Classification: A Systematic Review
por: van Gelderen, Lisanne, et al.
Publicado: (2024)
por: van Gelderen, Lisanne, et al.
Publicado: (2024)
Learning Disentangled Speech Representations
por: Brima, Yusuf, et al.
Publicado: (2023)
por: Brima, Yusuf, et al.
Publicado: (2023)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
por: Dutta, Soumya, et al.
Publicado: (2023)
por: Dutta, Soumya, et al.
Publicado: (2023)
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
por: Zhu, Haina, et al.
Publicado: (2025)
por: Zhu, Haina, et al.
Publicado: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
por: Liu, Zhijun, et al.
Publicado: (2024)
por: Liu, Zhijun, et al.
Publicado: (2024)
Real-time Speech Summarization for Medical Conversations
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
On the Role of Speech Data in Reducing Toxicity Detection Bias
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Scaling Rich Style-Prompted Text-to-Speech Datasets
por: Diwan, Anuj, et al.
Publicado: (2025)
por: Diwan, Anuj, et al.
Publicado: (2025)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
por: Dutta, Soumya, et al.
Publicado: (2025)
por: Dutta, Soumya, et al.
Publicado: (2025)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
por: Ghosh, Sreyan, et al.
Publicado: (2026)
por: Ghosh, Sreyan, et al.
Publicado: (2026)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
por: Bataev, Vladimir
Publicado: (2025)
por: Bataev, Vladimir
Publicado: (2025)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
por: Ok, Hyunjong, et al.
Publicado: (2025)
por: Ok, Hyunjong, et al.
Publicado: (2025)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
por: Jia, Dongya, et al.
Publicado: (2025)
por: Jia, Dongya, et al.
Publicado: (2025)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
por: Biswas, Subrata, et al.
Publicado: (2025)
por: Biswas, Subrata, et al.
Publicado: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization
por: Samin, Md. Nazmus Sadat, et al.
Publicado: (2024)
por: Samin, Md. Nazmus Sadat, et al.
Publicado: (2024)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
por: Li, Zehua Kcriss, et al.
Publicado: (2024)
por: Li, Zehua Kcriss, et al.
Publicado: (2024)
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
por: R, Aron, et al.
Publicado: (2024)
por: R, Aron, et al.
Publicado: (2024)
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
por: Xie, Tianxin, et al.
Publicado: (2024)
por: Xie, Tianxin, et al.
Publicado: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
por: Weise, Tobias, et al.
Publicado: (2024)
por: Weise, Tobias, et al.
Publicado: (2024)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
por: Inoue, Nakamasa, et al.
Publicado: (2024)
por: Inoue, Nakamasa, et al.
Publicado: (2024)
Ejemplares similares
-
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
por: Bhattacharya, Debarpan, et al.
Publicado: (2025) -
Improving Self-supervised Pre-training using Accent-Specific Codebooks
por: Prabhu, Darshan, et al.
Publicado: (2024) -
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
por: Dutta, Soumya, et al.
Publicado: (2024) -
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
por: Lian, Jiachen, et al.
Publicado: (2022) -
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
por: Mu, Zhaoxi, et al.
Publicado: (2023)