How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917695200428032 |
|---|---|
| author | Mousavi, Pooneh Duret, Jarod Zaiem, Salah Della Libera, Luca Ploujnikov, Artem Subakan, Cem Ravanelli, Mirco |
| author_facet | Mousavi, Pooneh Duret, Jarod Zaiem, Salah Della Libera, Luca Ploujnikov, Artem Subakan, Cem Ravanelli, Mirco |
| contents | Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_10735 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | How Should We Extract Discrete Audio Tokens from Self-Supervised Models? Mousavi, Pooneh Duret, Jarod Zaiem, Salah Della Libera, Luca Ploujnikov, Artem Subakan, Cem Ravanelli, Mirco Sound Artificial Intelligence Computation and Language Audio and Speech Processing Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications. |
| title | How Should We Extract Discrete Audio Tokens from Self-Supervised Models? |
| topic | Sound Artificial Intelligence Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.10735 |