On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Zijian, Zhou, Wei, Schlüter, Ralf, Ney, Hermann |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Label-Context-Dependent Internal Language Model Estimation for CTC
di: Yang, Zijian, et al.
Pubblicazione: (2025)
di: Yang, Zijian, et al.
Pubblicazione: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
di: Yang, Zijian, et al.
Pubblicazione: (2026)
di: Yang, Zijian, et al.
Pubblicazione: (2026)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
di: Raissi, Tina, et al.
Pubblicazione: (2025)
di: Raissi, Tina, et al.
Pubblicazione: (2025)
Unified Learnable 2D Convolutional Feature Extraction for ASR
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
di: Raissi, Tina, et al.
Pubblicazione: (2024)
di: Raissi, Tina, et al.
Pubblicazione: (2024)
The Conformer Encoder May Reverse the Time Dimension
di: Schmitt, Robin, et al.
Pubblicazione: (2024)
di: Schmitt, Robin, et al.
Pubblicazione: (2024)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
di: Wang, Peng, et al.
Pubblicazione: (2023)
di: Wang, Peng, et al.
Pubblicazione: (2023)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
Self-Supervised Learning for Multi-Channel Neural Transducer
di: Kojima, Atsushi
Pubblicazione: (2024)
di: Kojima, Atsushi
Pubblicazione: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Lightweight Transducer Based on Frame-Level Criterion
di: Wan, Genshun, et al.
Pubblicazione: (2024)
di: Wan, Genshun, et al.
Pubblicazione: (2024)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
di: Hilmes, Benedikt, et al.
Pubblicazione: (2025)
di: Hilmes, Benedikt, et al.
Pubblicazione: (2025)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
di: Tang, Yun, et al.
Pubblicazione: (2025)
di: Tang, Yun, et al.
Pubblicazione: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
di: Vieting, Peter, et al.
Pubblicazione: (2023)
di: Vieting, Peter, et al.
Pubblicazione: (2023)
Error Analysis in a Modular Meeting Transcription System
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
di: Gong, Xun, et al.
Pubblicazione: (2024)
di: Gong, Xun, et al.
Pubblicazione: (2024)
DDTSE: Discriminative Diffusion Model for Target Speech Extraction
di: Zhang, Leying, et al.
Pubblicazione: (2023)
di: Zhang, Leying, et al.
Pubblicazione: (2023)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
di: Sasu, David, et al.
Pubblicazione: (2025)
di: Sasu, David, et al.
Pubblicazione: (2025)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Medical Spoken Named Entity Recognition
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
Label-Looping: Highly Efficient Decoding for Transducers
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
di: Ko, Myeongjin, et al.
Pubblicazione: (2023)
di: Ko, Myeongjin, et al.
Pubblicazione: (2023)
Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
di: Rutowski, Tomek, et al.
Pubblicazione: (2024)
di: Rutowski, Tomek, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Label-Context-Dependent Internal Language Model Estimation for CTC
di: Yang, Zijian, et al.
Pubblicazione: (2025) -
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
di: Yang, Zijian, et al.
Pubblicazione: (2026) -
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
di: Raissi, Tina, et al.
Pubblicazione: (2025) -
Unified Learnable 2D Convolutional Feature Extraction for ASR
di: Vieting, Peter, et al.
Pubblicazione: (2025) -
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)