Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dhakal, Manish, Chhetri, Arman, Gupta, Aman Kumar, Lamichhane, Prabin, Pandey, Suraj, Shakya, Subarna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GMM-ResNet2: Ensemble of Group ResNet Networks for Synthetic Speech Detection
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
Squeeze-and-Excite ResNet-Conformers for Sound Event Localization, Detection, and Distance Estimation for DCASE 2024 Challenge
von: Yeow, Jun Wei, et al.
Veröffentlicht: (2024)
von: Yeow, Jun Wei, et al.
Veröffentlicht: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
von: Ahmad, Rehan, et al.
Veröffentlicht: (2026)
von: Ahmad, Rehan, et al.
Veröffentlicht: (2026)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
BabAR: from phoneme recognition to developmental measures of young children's speech production
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
von: Sereda, Taras
Veröffentlicht: (2024)
von: Sereda, Taras
Veröffentlicht: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
von: Ghimire, Rupak Raj, et al.
Veröffentlicht: (2024)
von: Ghimire, Rupak Raj, et al.
Veröffentlicht: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
von: Lee, Dongheon, et al.
Veröffentlicht: (2026)
von: Lee, Dongheon, et al.
Veröffentlicht: (2026)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Inter-channel Conv-TasNet for multichannel speech enhancement
von: Lee, Dongheon, et al.
Veröffentlicht: (2021)
von: Lee, Dongheon, et al.
Veröffentlicht: (2021)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
von: Ouyang, Qianhe
Veröffentlicht: (2025)
von: Ouyang, Qianhe
Veröffentlicht: (2025)
AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
von: Pieper, Jaden, et al.
Veröffentlicht: (2024)
von: Pieper, Jaden, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GMM-ResNet2: Ensemble of Group ResNet Networks for Synthetic Speech Detection
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024) -
Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024) -
Squeeze-and-Excite ResNet-Conformers for Sound Event Localization, Detection, and Distance Estimation for DCASE 2024 Challenge
von: Yeow, Jun Wei, et al.
Veröffentlicht: (2024) -
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025) -
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
von: Tang, Duowei, et al.
Veröffentlicht: (2023)