Enhancing CTC-based speech recognition with diverse modeling units
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Shiyi, Lei, Zhihong, Xu, Mingbin, Na, Xingyu, Huang, Zhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
SPMamba: State-space model is all you need in speech separation
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
A correlation-permutation approach for speech-music encoders model merging
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
Ensemble of classifiers for speech evaluation
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Online neural fusion of distortionless differential beamformers for robust speech enhancement
von: Qian, Yuanhang, et al.
Veröffentlicht: (2025)
von: Qian, Yuanhang, et al.
Veröffentlicht: (2025)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
von: Wan, Xucheng, et al.
Veröffentlicht: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
FINALLY: fast and universal speech enhancement with studio-like quality
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Selective Classifier-free Guidance for Zero-shot Text-to-speech
von: Zheng, John, et al.
Veröffentlicht: (2025)
von: Zheng, John, et al.
Veröffentlicht: (2025)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
Semantic visually-guided acoustic highlighting with large vision-language models
von: Huang, Junhua, et al.
Veröffentlicht: (2026)
von: Huang, Junhua, et al.
Veröffentlicht: (2026)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
Towards measuring fairness in speech recognition: Fair-Speech dataset
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on Heart Sounds
von: Han, Jiho, et al.
Veröffentlicht: (2025)
von: Han, Jiho, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026) -
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024) -
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024) -
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
von: Lei, Zhihong, et al.
Veröffentlicht: (2024) -
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)