Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Yadav, Hemant, Sitaram, Sunayana, Shah, Rajiv Ratn |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
por: Sahipjohn, Neha, et al.
Publicado: (2024)
por: Sahipjohn, Neha, et al.
Publicado: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
por: Wu, Wenxuan, et al.
Publicado: (2024)
por: Wu, Wenxuan, et al.
Publicado: (2024)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
por: Wang, Tianrui, et al.
Publicado: (2024)
por: Wang, Tianrui, et al.
Publicado: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
por: Biyani, Ishan D., et al.
Publicado: (2025)
por: Biyani, Ishan D., et al.
Publicado: (2025)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2024)
por: Peng, Junyi, et al.
Publicado: (2024)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
por: Whetten, Ryan, et al.
Publicado: (2026)
por: Whetten, Ryan, et al.
Publicado: (2026)
Selecting N-lowest scores for training MOS prediction models
por: Kondo, Yuto, et al.
Publicado: (2025)
por: Kondo, Yuto, et al.
Publicado: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
por: Feng, Tiantian, et al.
Publicado: (2023)
por: Feng, Tiantian, et al.
Publicado: (2023)
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
por: Sun, Xingwei, et al.
Publicado: (2025)
por: Sun, Xingwei, et al.
Publicado: (2025)
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
por: Feng, Jianyuan, et al.
Publicado: (2025)
por: Feng, Jianyuan, et al.
Publicado: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
por: Zhang, Yizhou, et al.
Publicado: (2025)
por: Zhang, Yizhou, et al.
Publicado: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
por: Liu, Alexander H., et al.
Publicado: (2025)
por: Liu, Alexander H., et al.
Publicado: (2025)
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
por: Roy, Arnab Kumar, et al.
Publicado: (2025)
por: Roy, Arnab Kumar, et al.
Publicado: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
por: Niizumi, Daisuke, et al.
Publicado: (2024)
por: Niizumi, Daisuke, et al.
Publicado: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)
por: Pham, The Hieu, et al.
Publicado: (2025)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
por: Wang, Huimeng, et al.
Publicado: (2024)
por: Wang, Huimeng, et al.
Publicado: (2024)
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
por: Zhang, Yixiao, et al.
Publicado: (2025)
por: Zhang, Yixiao, et al.
Publicado: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
por: Zhu, Yongxin, et al.
Publicado: (2024)
por: Zhu, Yongxin, et al.
Publicado: (2024)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
por: Wang, Linqin, et al.
Publicado: (2024)
por: Wang, Linqin, et al.
Publicado: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
por: Zhang, Leying, et al.
Publicado: (2025)
por: Zhang, Leying, et al.
Publicado: (2025)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
por: Wang, Yuancheng, et al.
Publicado: (2025)
por: Wang, Yuancheng, et al.
Publicado: (2025)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
por: Yadav, Hemant, et al.
Publicado: (2024)
por: Yadav, Hemant, et al.
Publicado: (2024)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
por: Kheir, Yassine El, et al.
Publicado: (2024)
por: Kheir, Yassine El, et al.
Publicado: (2024)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
por: Poncelet, Jakob, et al.
Publicado: (2023)
por: Poncelet, Jakob, et al.
Publicado: (2023)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
por: Sang, Mufan, et al.
Publicado: (2024)
por: Sang, Mufan, et al.
Publicado: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
por: Zhang, Jisi, et al.
Publicado: (2024)
por: Zhang, Jisi, et al.
Publicado: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
por: Ma, Ding, et al.
Publicado: (2026)
por: Ma, Ding, et al.
Publicado: (2026)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
por: Cai, Pengfei, et al.
Publicado: (2024)
por: Cai, Pengfei, et al.
Publicado: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
por: Zeng, Aohan, et al.
Publicado: (2024)
por: Zeng, Aohan, et al.
Publicado: (2024)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
por: Dutta, Soumya, et al.
Publicado: (2025)
por: Dutta, Soumya, et al.
Publicado: (2025)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
por: Liu, Xiaoyang, et al.
Publicado: (2025)
por: Liu, Xiaoyang, et al.
Publicado: (2025)
Ejemplares similares
-
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
por: Sahipjohn, Neha, et al.
Publicado: (2024) -
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024) -
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
por: Wu, Wenxuan, et al.
Publicado: (2024) -
Progressive Residual Extraction based Pre-training for Speech Representation Learning
por: Wang, Tianrui, et al.
Publicado: (2024) -
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
por: Zhu, Xiaoxu, et al.
Publicado: (2025)