Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyeonseung, Yoon, Ji Won, Kim, Sungsoo, Kim, Nam Soo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
Multi-blank Transducers for Speech Recognition
by: Xu, Hainan, et al.
Published: (2022)
by: Xu, Hainan, et al.
Published: (2022)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022)
by: Yoon, Ji Won, et al.
Published: (2022)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
by: Lee, Joun Yeop, et al.
Published: (2024)
by: Lee, Joun Yeop, et al.
Published: (2024)
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
by: Kim, Semin, et al.
Published: (2024)
by: Kim, Semin, et al.
Published: (2024)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
by: Eom, SooHwan, et al.
Published: (2024)
by: Eom, SooHwan, et al.
Published: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
by: Xu, Hainan, et al.
Published: (2024)
by: Xu, Hainan, et al.
Published: (2024)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
by: Zeineldeen, Mohammad, et al.
Published: (2023)
by: Zeineldeen, Mohammad, et al.
Published: (2023)
FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
by: Ahn, Sunghwan, et al.
Published: (2025)
by: Ahn, Sunghwan, et al.
Published: (2025)
Efficient Parallel Audio Generation using Group Masked Language Modeling
by: Jeong, Myeonghun, et al.
Published: (2024)
by: Jeong, Myeonghun, et al.
Published: (2024)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
by: Kim, Hyeongju, et al.
Published: (2025)
by: Kim, Hyeongju, et al.
Published: (2025)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
by: Bataev, Vladimir
Published: (2025)
by: Bataev, Vladimir
Published: (2025)
Scalable Frameworks for Real-World Audio-Visual Speech Recognition
by: Kim, Sungnyun
Published: (2025)
by: Kim, Sungnyun
Published: (2025)
Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
by: Mansour, Mohamed F.
Published: (2024)
by: Mansour, Mohamed F.
Published: (2024)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
by: Yang, Jinhyeok, et al.
Published: (2026)
by: Yang, Jinhyeok, et al.
Published: (2026)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
by: Deng, Keqi, et al.
Published: (2024)
by: Deng, Keqi, et al.
Published: (2024)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
by: Yang, Zijian, et al.
Published: (2026)
by: Yang, Zijian, et al.
Published: (2026)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
by: Kim, Seungmin, et al.
Published: (2025)
by: Kim, Seungmin, et al.
Published: (2025)
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
by: Cui, Mingyu, et al.
Published: (2025)
by: Cui, Mingyu, et al.
Published: (2025)
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
by: Dmitrieva, Ekaterina, et al.
Published: (2025)
by: Dmitrieva, Ekaterina, et al.
Published: (2025)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
by: Feng, Chen, et al.
Published: (2025)
by: Feng, Chen, et al.
Published: (2025)
Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Towards Audio Codec-based Speech Separation
by: Yip, Jia Qi, et al.
Published: (2024)
by: Yip, Jia Qi, et al.
Published: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
by: Dhawan, Kunal, et al.
Published: (2024)
by: Dhawan, Kunal, et al.
Published: (2024)
EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
by: Thimonier, Hugo, et al.
Published: (2025)
by: Thimonier, Hugo, et al.
Published: (2025)
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical Diagnostics
by: Kumar, Kabir
Published: (2025)
by: Kumar, Kabir
Published: (2025)
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
by: Upadhyay, Shreya G., et al.
Published: (2024)
by: Upadhyay, Shreya G., et al.
Published: (2024)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
by: Zhou, Nanjun, et al.
Published: (2025)
by: Zhou, Nanjun, et al.
Published: (2025)
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
by: Upadhyay, Shreya G., et al.
Published: (2024)
by: Upadhyay, Shreya G., et al.
Published: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
by: Kim, Ji-Hoon, et al.
Published: (2023)
by: Kim, Ji-Hoon, et al.
Published: (2023)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
by: Jiang, Xilin, et al.
Published: (2024)
by: Jiang, Xilin, et al.
Published: (2024)
Adapting WavLM for Speech Emotion Recognition
by: Diatlova, Daria, et al.
Published: (2024)
by: Diatlova, Daria, et al.
Published: (2024)
Keyword-Guided Adaptation of Automatic Speech Recognition
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Test-Time Adaptation for Speech Emotion Recognition
by: Dong, Jiaheng, et al.
Published: (2026)
by: Dong, Jiaheng, et al.
Published: (2026)
Similar Items
-
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
by: Kim, Minchan, et al.
Published: (2024) -
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
by: Kim, Minchan, et al.
Published: (2024) -
Multi-blank Transducers for Speech Recognition
by: Xu, Hainan, et al.
Published: (2022) -
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022) -
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
by: Lee, Joun Yeop, et al.
Published: (2024)