Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Jiamin, Li, Ke, Guo, Jinxi, Tjandra, Andros, Shangguan, Yuan, Sari, Leda, Wu, Chunyang, Jia, Junteng, Mahadeokar, Jay, Kalinli, Ozlem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
Towards scalable efficient on-device ASR with transfer learning
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
von: Irigoyen, Julian, et al.
Veröffentlicht: (2025)
von: Irigoyen, Julian, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Index-ASR Technical Report
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
Advocating Character Error Rate for Multilingual ASR Evaluation
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
Efficient Multilingual ASR Finetuning via LoRA Language Experts
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
von: Huang, W. Ronny, et al.
Veröffentlicht: (2024)
von: Huang, W. Ronny, et al.
Veröffentlicht: (2024)
persoDA: Personalized Data Augmentation for Personalized ASR
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Efficient Adaptation of Multilingual Models for Japanese ASR
von: Bajo, Mark, et al.
Veröffentlicht: (2024)
von: Bajo, Mark, et al.
Veröffentlicht: (2024)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
von: Yan, Brian, et al.
Veröffentlicht: (2024)
von: Yan, Brian, et al.
Veröffentlicht: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024) -
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024) -
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024) -
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024) -
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)