Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Caubrière, Antoine, Gauthier, Elodie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
von: Beyene, Luel Hagos, et al.
Veröffentlicht: (2025)
von: Beyene, Luel Hagos, et al.
Veröffentlicht: (2025)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
How Redundant Is the Transformer Stack in Speech Representation Models?
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
Improving Speech Decoding from ECoG with Self-Supervised Pretraining
von: Yuan, Brian A., et al.
Veröffentlicht: (2024)
von: Yuan, Brian A., et al.
Veröffentlicht: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022) -
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024) -
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024) -
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026) -
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)