Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models
Fuente:
arXiv
Saved in:
| Main Authors: | Haider, Adnan, Na, Xingyu, McDermott, Erik, Ng, Tim, Huang, Zhen, Zhuang, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
by: Hori, Takaaki, et al.
Published: (2025)
by: Hori, Takaaki, et al.
Published: (2025)
ACES: Automatic Cohort Extraction System for Event-Stream Datasets
by: Xu, Justin, et al.
Published: (2024)
by: Xu, Justin, et al.
Published: (2024)
Embedding Compression for Efficient Re-Identification
by: McDermott, Luke
Published: (2024)
by: McDermott, Luke
Published: (2024)
Finding Stable Subnetworks at Initialization with Dataset Distillation
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)
by: McDermott, Luke, et al.
Published: (2023)
Text Conditioned Symbolic Drumbeat Generation using Latent Diffusion Models
by: Jajoria, Pushkar, et al.
Published: (2024)
by: Jajoria, Pushkar, et al.
Published: (2024)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
by: Lee, Hyeonseung, et al.
Published: (2024)
by: Lee, Hyeonseung, et al.
Published: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
by: Saif, A F M, et al.
Published: (2024)
by: Saif, A F M, et al.
Published: (2024)
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
by: Rossenbach, Nick, et al.
Published: (2025)
by: Rossenbach, Nick, et al.
Published: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
by: Dhawan, Kunal, et al.
Published: (2024)
by: Dhawan, Kunal, et al.
Published: (2024)
Higher-Order Singular-Value Derivatives of Rectangular Real Matrices
by: Luo, Róisín, et al.
Published: (2025)
by: Luo, Róisín, et al.
Published: (2025)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
by: Bartelds, Martijn, et al.
Published: (2025)
by: Bartelds, Martijn, et al.
Published: (2025)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
by: Rufai, Amina Mardiyyah, et al.
Published: (2020)
by: Rufai, Amina Mardiyyah, et al.
Published: (2020)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
by: Hilmes, Benedikt, et al.
Published: (2024)
by: Hilmes, Benedikt, et al.
Published: (2024)
Revisiting ASR Error Correction with Specialized Models
by: Gu, Zijin, et al.
Published: (2024)
by: Gu, Zijin, et al.
Published: (2024)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
by: Xu, Mingbin, et al.
Published: (2023)
by: Xu, Mingbin, et al.
Published: (2023)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
Training for Speech Recognition on Coprocessors
by: Baunsgaard, Sebastian, et al.
Published: (2020)
by: Baunsgaard, Sebastian, et al.
Published: (2020)
ADAPTive Input Training for Many-to-One Pre-Training on Time-Series Classification
by: Quinlan, Paul, et al.
Published: (2026)
by: Quinlan, Paul, et al.
Published: (2026)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
by: Hono, Yukiya, et al.
Published: (2023)
by: Hono, Yukiya, et al.
Published: (2023)
Enhancing Speech Emotion Recognition via Fine-Tuning Pre-Trained Models and Hyper-Parameter Optimisation
by: Golbaghi, Aryan, et al.
Published: (2025)
by: Golbaghi, Aryan, et al.
Published: (2025)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
by: Ngo, Huong, et al.
Published: (2025)
by: Ngo, Huong, et al.
Published: (2025)
Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Training More Robust Classification Model via Discriminative Loss and Gaussian Noise Injection
by: Nguyen, Hai-Vy, et al.
Published: (2024)
by: Nguyen, Hai-Vy, et al.
Published: (2024)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Optimization-Induced Dynamics of Lipschitz Continuity in Neural Networks
by: Luo, Róisín, et al.
Published: (2025)
by: Luo, Róisín, et al.
Published: (2025)
AutoHete: An Automatic and Efficient Heterogeneous Training System for LLMs
by: Zeng, Zihao, et al.
Published: (2025)
by: Zeng, Zihao, et al.
Published: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
by: Yang, Zijian, et al.
Published: (2026)
by: Yang, Zijian, et al.
Published: (2026)
Optimizing Byte-level Representation for End-to-end ASR
by: Hsiao, Roger, et al.
Published: (2024)
by: Hsiao, Roger, et al.
Published: (2024)
Keyword-Guided Adaptation of Automatic Speech Recognition
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
by: Galvez, Daniel, et al.
Published: (2024)
by: Galvez, Daniel, et al.
Published: (2024)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
by: Zeineldeen, Mohammad, et al.
Published: (2023)
by: Zeineldeen, Mohammad, et al.
Published: (2023)
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
by: Cheng, Hailing, et al.
Published: (2026)
by: Cheng, Hailing, et al.
Published: (2026)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
by: Wang, Charles L., et al.
Published: (2026)
by: Wang, Charles L., et al.
Published: (2026)
Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation
by: Yin, Xuwang, et al.
Published: (2025)
by: Yin, Xuwang, et al.
Published: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
by: Yao, Zengwei, et al.
Published: (2024)
by: Yao, Zengwei, et al.
Published: (2024)
Accelerating Recommender Model Training by Dynamically Skipping Stale Embeddings
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
Similar Items
-
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
by: Hori, Takaaki, et al.
Published: (2025) -
ACES: Automatic Cohort Extraction System for Event-Stream Datasets
by: Xu, Justin, et al.
Published: (2024) -
Embedding Compression for Efficient Re-Identification
by: McDermott, Luke
Published: (2024) -
Finding Stable Subnetworks at Initialization with Dataset Distillation
by: McDermott, Luke, et al.
Published: (2025) -
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)