Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhaoqing, Xu, Haoning, Jin, Zengrui, Meng, Lingwei, Wang, Tianzi, Wang, Huimeng, Chen, Youjun, Cui, Mingyu, Hu, Shujie, Liu, Xunying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
von: Wang, Tianzi, et al.
Veröffentlicht: (2024)
von: Wang, Tianzi, et al.
Veröffentlicht: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
von: Geng, Mengzhe, et al.
Veröffentlicht: (2024)
von: Geng, Mengzhe, et al.
Veröffentlicht: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Multi-bit Audio Watermarking
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Complexity boosted adaptive training for better low resource ASR performance
von: Lu, Hongxuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongxuan, et al.
Veröffentlicht: (2024)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition
von: Deng, Chengxi, et al.
Veröffentlicht: (2025)
von: Deng, Chengxi, et al.
Veröffentlicht: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Index-ASR Technical Report
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024) -
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024) -
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025) -
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025) -
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)