Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | San, Nay, Paraskevopoulos, Georgios, Arora, Aryaman, He, Xiluo, Kaur, Prabhjot, Adams, Oliver, Jurafsky, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
Direct Punjabi to English speech translation using discrete units
di: Kaur, Prabhjot, et al.
Pubblicazione: (2024)
di: Kaur, Prabhjot, et al.
Pubblicazione: (2024)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
di: Tang, Duowei, et al.
Pubblicazione: (2023)
di: Tang, Duowei, et al.
Pubblicazione: (2023)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
di: Sereda, Taras
Pubblicazione: (2024)
di: Sereda, Taras
Pubblicazione: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
di: Damianos, Dimitrios, et al.
Pubblicazione: (2025)
di: Damianos, Dimitrios, et al.
Pubblicazione: (2025)
TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2026)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2026)
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
Prominence-aware automatic speech recognition for conversational speech
di: Linke, Julian, et al.
Pubblicazione: (2025)
di: Linke, Julian, et al.
Pubblicazione: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
di: Ahmad, Rehan, et al.
Pubblicazione: (2026)
di: Ahmad, Rehan, et al.
Pubblicazione: (2026)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
di: Fang, Yangui, et al.
Pubblicazione: (2025)
di: Fang, Yangui, et al.
Pubblicazione: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
di: Wright, George August, et al.
Pubblicazione: (2023)
di: Wright, George August, et al.
Pubblicazione: (2023)
Automated evaluation of children's speech fluency for low-resource languages
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
di: Damianos, Dimitrios, et al.
Pubblicazione: (2025)
di: Damianos, Dimitrios, et al.
Pubblicazione: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
di: He, Xiluo, et al.
Pubblicazione: (2025)
di: He, Xiluo, et al.
Pubblicazione: (2025)
BabAR: from phoneme recognition to developmental measures of young children's speech production
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
di: Gogoi, Parismita, et al.
Pubblicazione: (2025)
di: Gogoi, Parismita, et al.
Pubblicazione: (2025)
Language model integration based on memory control for sequence to sequence speech recognition
di: Cho, Jaejin, et al.
Pubblicazione: (2018)
di: Cho, Jaejin, et al.
Pubblicazione: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
Guiding the underwater acoustic target recognition with interpretable contrastive learning
di: Xie, Yuan, et al.
Pubblicazione: (2024)
di: Xie, Yuan, et al.
Pubblicazione: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
di: Lonergan, Liam, et al.
Pubblicazione: (2024)
di: Lonergan, Liam, et al.
Pubblicazione: (2024)
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer
di: Kim, Jeong-Yoon, et al.
Pubblicazione: (2024)
di: Kim, Jeong-Yoon, et al.
Pubblicazione: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
di: Pandey, Rahul, et al.
Pubblicazione: (2023)
di: Pandey, Rahul, et al.
Pubblicazione: (2023)
Monaural speech enhancement on drone via Adapter based transfer learning
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
di: Gong, Rong, et al.
Pubblicazione: (2024)
di: Gong, Rong, et al.
Pubblicazione: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024) -
Direct Punjabi to English speech translation using discrete units
di: Kaur, Prabhjot, et al.
Pubblicazione: (2024) -
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
di: Tang, Duowei, et al.
Pubblicazione: (2023) -
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
di: Sereda, Taras
Pubblicazione: (2024) -
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)