Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.19612 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909973714305024 |
|---|---|
| author | Tandazo, Angelo Ortiz Khentout, Manel Benchekroun, Youssef Hueber, Thomas Dupoux, Emmanuel |
| author_facet | Tandazo, Angelo Ortiz Khentout, Manel Benchekroun, Youssef Hueber, Thomas Dupoux, Emmanuel |
| contents | This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_19612 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery Tandazo, Angelo Ortiz Khentout, Manel Benchekroun, Youssef Hueber, Thomas Dupoux, Emmanuel Computation and Language Audio and Speech Processing This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models. |
| title | MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.19612 |