Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | San, Nay, Paraskevopoulos, Georgios, Arora, Aryaman, He, Xiluo, Kaur, Prabhjot, Adams, Oliver, Jurafsky, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
von: Sereda, Taras
Veröffentlicht: (2024)
von: Sereda, Taras
Veröffentlicht: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2026)
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2026)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
von: Ahmad, Rehan, et al.
Veröffentlicht: (2026)
von: Ahmad, Rehan, et al.
Veröffentlicht: (2026)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
BabAR: from phoneme recognition to developmental measures of young children's speech production
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
Guiding the underwater acoustic target recognition with interpretable contrastive learning
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
von: Lonergan, Liam, et al.
Veröffentlicht: (2024)
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024) -
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024) -
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
von: Tang, Duowei, et al.
Veröffentlicht: (2023) -
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
von: Sereda, Taras
Veröffentlicht: (2024) -
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)