Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Hemant, Sitaram, Sunayana, Shah, Rajiv Ratn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
von: Sun, Xingwei, et al.
Veröffentlicht: (2025)
von: Sun, Xingwei, et al.
Veröffentlicht: (2025)
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
von: Feng, Jianyuan, et al.
Veröffentlicht: (2025)
von: Feng, Jianyuan, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
von: Roy, Arnab Kumar, et al.
Veröffentlicht: (2025)
von: Roy, Arnab Kumar, et al.
Veröffentlicht: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
von: Zhang, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
von: Wang, Linqin, et al.
Veröffentlicht: (2024)
von: Wang, Linqin, et al.
Veröffentlicht: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2024)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
von: Poncelet, Jakob, et al.
Veröffentlicht: (2023)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2023)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024) -
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024) -
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024) -
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024) -
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)