NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Minglun, Bai, Ye, Shen, Chen, Huang, Youjia, Huang, Mingkun, Lin, Zehua, Dong, Linhao, Lu, Lu, Wang, Yuxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
von: Poncelet, Jakob, et al.
Veröffentlicht: (2021)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2021)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
von: Ma, Duo, et al.
Veröffentlicht: (2024)
von: Ma, Duo, et al.
Veröffentlicht: (2024)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
Next Tokens Denoising for Speech Synthesis
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Rethinking Mamba in Speech Processing by Self-Supervised Models
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
von: Chung, Raymond
Veröffentlicht: (2026)
von: Chung, Raymond
Veröffentlicht: (2026)
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2025)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2025)
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
von: Li, Zehan, et al.
Veröffentlicht: (2025)
von: Li, Zehan, et al.
Veröffentlicht: (2025)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
von: Hanilçi, Cemal, et al.
Veröffentlicht: (2026)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
von: Li, Ruiqi, et al.
Veröffentlicht: (2024) -
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025) -
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025) -
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
von: Huang, He, et al.
Veröffentlicht: (2024) -
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)