Exploring SSL Discrete Tokens for Multilingual ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Mingyu, Tan, Daxin, Yang, Yifan, Wang, Dingdong, Wang, Huimeng, Chen, Xiao, Chen, Xie, Liu, Xunying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
von: Cui, Wenqian, et al.
Veröffentlicht: (2026)
von: Cui, Wenqian, et al.
Veröffentlicht: (2026)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
von: Huang, W. Ronny, et al.
Veröffentlicht: (2024)
von: Huang, W. Ronny, et al.
Veröffentlicht: (2024)
Advocating Character Error Rate for Multilingual ASR Evaluation
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Efficient Multilingual ASR Finetuning via LoRA Language Experts
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
von: Wang, Peng, et al.
Veröffentlicht: (2023)
von: Wang, Peng, et al.
Veröffentlicht: (2023)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
von: Yan, Brian, et al.
Veröffentlicht: (2024)
von: Yan, Brian, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024) -
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024) -
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025) -
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025) -
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)