How Should We Extract Discrete Audio Tokens from Self-Supervised Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mousavi, Pooneh, Duret, Jarod, Zaiem, Salah, Della Libera, Luca, Ploujnikov, Artem, Subakan, Cem, Ravanelli, Mirco
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917695200428032
author Mousavi, Pooneh
Duret, Jarod
Zaiem, Salah
Della Libera, Luca
Ploujnikov, Artem
Subakan, Cem
Ravanelli, Mirco
author_facet Mousavi, Pooneh
Duret, Jarod
Zaiem, Salah
Della Libera, Luca
Ploujnikov, Artem
Subakan, Cem
Ravanelli, Mirco
contents Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10735
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
Mousavi, Pooneh
Duret, Jarod
Zaiem, Salah
Della Libera, Luca
Ploujnikov, Artem
Subakan, Cem
Ravanelli, Mirco
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications.
title How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.10735