Scaling HuBERT for African Languages: From Base to Large and XL
Fuente:
arXiv
Saved in:
| Main Authors: | Caubrière, Antoine, Gauthier, Elodie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
by: Caubrière, Antoine, et al.
Published: (2024)
by: Caubrière, Antoine, et al.
Published: (2024)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022)
by: Yoon, Ji Won, et al.
Published: (2022)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
MelHuBERT: A simplified HuBERT on Mel spectrograms
by: Lin, Tzu-Quan, et al.
Published: (2022)
by: Lin, Tzu-Quan, et al.
Published: (2022)
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024)
by: Boito, Marcely Zanon, et al.
Published: (2024)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
by: Komatsu, Ryota, et al.
Published: (2024)
by: Komatsu, Ryota, et al.
Published: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
by: Li, Xingyuan, et al.
Published: (2024)
by: Li, Xingyuan, et al.
Published: (2024)
ExHuBERT: Enhancing HuBERT Through Block Extension and Fine-Tuning on 37 Emotion Datasets
by: Amiriparian, Shahin, et al.
Published: (2024)
by: Amiriparian, Shahin, et al.
Published: (2024)
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
by: Vielzeuf, Valentin
Published: (2024)
by: Vielzeuf, Valentin
Published: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
by: Papadopoulos, Aristeidis, et al.
Published: (2026)
by: Papadopoulos, Aristeidis, et al.
Published: (2026)
Artificial Rigidities vs. Biological Noise: A Comparative Analysis of Multisensory Integration in AV-HuBERT and Human Observers
by: López, Francisco Portillo
Published: (2026)
by: López, Francisco Portillo
Published: (2026)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
by: Huo, Robin, et al.
Published: (2025)
by: Huo, Robin, et al.
Published: (2025)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
by: Chi, Hyung Gun, et al.
Published: (2025)
by: Chi, Hyung Gun, et al.
Published: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal
by: Gauthier, Elodie, et al.
Published: (2024)
by: Gauthier, Elodie, et al.
Published: (2024)
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
by: Ghazal, Nizar El, et al.
Published: (2025)
by: Ghazal, Nizar El, et al.
Published: (2025)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
by: Yang, Haote, et al.
Published: (2025)
by: Yang, Haote, et al.
Published: (2025)
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
by: Türker, Melikşah, et al.
Published: (2025)
by: Türker, Melikşah, et al.
Published: (2025)
An investigation of structures responsible for gender bias in BERT and DistilBERT
by: Leteno, Thibaud, et al.
Published: (2024)
by: Leteno, Thibaud, et al.
Published: (2024)
Extending Translate-Train for ColBERT-X to African Language CLIR
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
HuRef: HUman-REadable Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2023)
by: Zeng, Boyi, et al.
Published: (2023)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
by: Huang, Zikang, et al.
Published: (2026)
by: Huang, Zikang, et al.
Published: (2026)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
by: Wu, Wenxuan, et al.
Published: (2024)
by: Wu, Wenxuan, et al.
Published: (2024)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
by: Chaffin, Antoine, et al.
Published: (2026)
by: Chaffin, Antoine, et al.
Published: (2026)
EuroBERT: Scaling Multilingual Encoders for European Languages
by: Boizard, Nicolas, et al.
Published: (2025)
by: Boizard, Nicolas, et al.
Published: (2025)
Analysis of Argument Structure Constructions in the Large Language Model BERT
by: Ramezani, Pegah, et al.
Published: (2024)
by: Ramezani, Pegah, et al.
Published: (2024)
PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
Detecting Bias in Large Language Models: Fine-tuned KcBERT
by: Lee, J. K., et al.
Published: (2024)
by: Lee, J. K., et al.
Published: (2024)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
by: Ismail, Saifelden M.
Published: (2025)
by: Ismail, Saifelden M.
Published: (2025)
DunbaaBERT: From Sacrifice to Semantics
by: Maab, Iffat, et al.
Published: (2026)
by: Maab, Iffat, et al.
Published: (2026)
SpikeBERT: A Language Spikformer Learned from BERT with Knowledge Distillation
by: Lv, Changze, et al.
Published: (2023)
by: Lv, Changze, et al.
Published: (2023)
Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support
by: Wu, Xiaojun, et al.
Published: (2024)
by: Wu, Xiaojun, et al.
Published: (2024)
From BERT to LLMs: Comparing and Understanding Chinese Classifier Prediction in Language Models
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
EgyBERT: A Large Language Model Pretrained on Egyptian Dialect Corpora
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
CultureBERT: Measuring Corporate Culture With Transformer-Based Language Models
by: Koch, Sebastian, et al.
Published: (2022)
by: Koch, Sebastian, et al.
Published: (2022)
MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
by: Yuan, Junming, et al.
Published: (2025)
by: Yuan, Junming, et al.
Published: (2025)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
by: Ma, Duo, et al.
Published: (2024)
by: Ma, Duo, et al.
Published: (2024)
Similar Items
-
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
by: Caubrière, Antoine, et al.
Published: (2024) -
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022) -
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
by: Cho, Cheol Jun, et al.
Published: (2023) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
by: Lin, Tzu-Quan, et al.
Published: (2022) -
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024)