Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gogoi, Parismita, Kalita, Sishir, Lalhminghlui, Wendy, Terhiija, Viyazonuo, Tzudir, Moakala, Sarmah, Priyankoo, Prasanna, S. R. M. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring rhythm formant analysis for Indic language classification
par: Gogoi, Parismita, et autres
Publié: (2024)
par: Gogoi, Parismita, et autres
Publié: (2024)
Analyzing long-term rhythm variations in Mising and Assamese using frequency domain correlates
par: Gogoi, Parismita, et autres
Publié: (2024)
par: Gogoi, Parismita, et autres
Publié: (2024)
Towards Prosodically Informed Mizo TTS without Explicit Tone Markings
par: Mohanta, Abhijit, et autres
Publié: (2026)
par: Mohanta, Abhijit, et autres
Publié: (2026)
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
par: Gogoi, Parismita, et autres
Publié: (2025)
par: Gogoi, Parismita, et autres
Publié: (2025)
Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh
par: Gogoi, Deepshikha, et autres
Publié: (2026)
par: Gogoi, Deepshikha, et autres
Publié: (2026)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
par: Sereda, Taras
Publié: (2024)
par: Sereda, Taras
Publié: (2024)
How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer
par: Kim, Minu, et autres
Publié: (2025)
par: Kim, Minu, et autres
Publié: (2025)
Automated evaluation of children's speech fluency for low-resource languages
par: Zhang, Bowen, et autres
Publié: (2025)
par: Zhang, Bowen, et autres
Publié: (2025)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
par: San, Nay, et autres
Publié: (2024)
par: San, Nay, et autres
Publié: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
par: Ma, Te, et autres
Publié: (2025)
par: Ma, Te, et autres
Publié: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
par: Paraskevopoulos, Georgios, et autres
Publié: (2024)
par: Paraskevopoulos, Georgios, et autres
Publié: (2024)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
par: Tang, Duowei, et autres
Publié: (2023)
par: Tang, Duowei, et autres
Publié: (2023)
Prominence-aware automatic speech recognition for conversational speech
par: Linke, Julian, et autres
Publié: (2025)
par: Linke, Julian, et autres
Publié: (2025)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
par: Dhakal, Manish, et autres
Publié: (2024)
par: Dhakal, Manish, et autres
Publié: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
par: Wright, George August, et autres
Publié: (2023)
par: Wright, George August, et autres
Publié: (2023)
Introduction to speech recognition
par: Dauphin, Gabriel
Publié: (2024)
par: Dauphin, Gabriel
Publié: (2024)
Improving child speech recognition with augmented child-like speech
par: Zhang, Yuanyuan, et autres
Publié: (2024)
par: Zhang, Yuanyuan, et autres
Publié: (2024)
Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
par: N, Rishith Sadashiv T, et autres
Publié: (2025)
par: N, Rishith Sadashiv T, et autres
Publié: (2025)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
par: Lonergan, Liam, et autres
Publié: (2024)
par: Lonergan, Liam, et autres
Publié: (2024)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
par: Akkiraju, Bhavana, et autres
Publié: (2025)
par: Akkiraju, Bhavana, et autres
Publié: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
par: Ducorroy, Alexandre, et autres
Publié: (2025)
par: Ducorroy, Alexandre, et autres
Publié: (2025)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
par: Vecino, Biel Tura, et autres
Publié: (2025)
par: Vecino, Biel Tura, et autres
Publié: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
par: Hu, Yuxuan, et autres
Publié: (2025)
par: Hu, Yuxuan, et autres
Publié: (2025)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
par: Ahmad, Rehan, et autres
Publié: (2026)
par: Ahmad, Rehan, et autres
Publié: (2026)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
par: Fang, Yangui, et autres
Publié: (2025)
par: Fang, Yangui, et autres
Publié: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
par: Liu, Xueyu, et autres
Publié: (2024)
par: Liu, Xueyu, et autres
Publié: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
par: Luo, Jiachen, et autres
Publié: (2025)
par: Luo, Jiachen, et autres
Publié: (2025)
BabAR: from phoneme recognition to developmental measures of young children's speech production
par: Lavechin, Marvin, et autres
Publié: (2026)
par: Lavechin, Marvin, et autres
Publié: (2026)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
par: Gong, Rong, et autres
Publié: (2024)
par: Gong, Rong, et autres
Publié: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
par: Cho, Jaejin, et autres
Publié: (2018)
par: Cho, Jaejin, et autres
Publié: (2018)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
par: Kesiraju, Santosh, et autres
Publié: (2023)
par: Kesiraju, Santosh, et autres
Publié: (2023)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
par: Dong, Lukuang, et autres
Publié: (2026)
par: Dong, Lukuang, et autres
Publié: (2026)
Zipformer: A faster and better encoder for automatic speech recognition
par: Yao, Zengwei, et autres
Publié: (2023)
par: Yao, Zengwei, et autres
Publié: (2023)
Self-consistent context aware conformer transducer for speech recognition
par: Kolokolov, Konstantin, et autres
Publié: (2024)
par: Kolokolov, Konstantin, et autres
Publié: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
par: Ma, Te, et autres
Publié: (2025)
par: Ma, Te, et autres
Publié: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
par: Han, Shiyi, et autres
Publié: (2024)
par: Han, Shiyi, et autres
Publié: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
par: Zheng, Naijun, et autres
Publié: (2024)
par: Zheng, Naijun, et autres
Publié: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
par: Yao, Zengwei, et autres
Publié: (2024)
par: Yao, Zengwei, et autres
Publié: (2024)
Robustifying automatic speech recognition by extracting slowly varying features
par: Pizarro, Matías, et autres
Publié: (2021)
par: Pizarro, Matías, et autres
Publié: (2021)
Documents similaires
-
Exploring rhythm formant analysis for Indic language classification
par: Gogoi, Parismita, et autres
Publié: (2024) -
Analyzing long-term rhythm variations in Mising and Assamese using frequency domain correlates
par: Gogoi, Parismita, et autres
Publié: (2024) -
Towards Prosodically Informed Mizo TTS without Explicit Tone Markings
par: Mohanta, Abhijit, et autres
Publié: (2026) -
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
par: Gogoi, Parismita, et autres
Publié: (2025) -
Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh
par: Gogoi, Deepshikha, et autres
Publié: (2026)