USAD: Universal Speech and Audio Representation via Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908492584976384 |
|---|---|
| author | Chang, Heng-Jui Bhati, Saurabhchand Glass, James Liu, Alexander H. |
| author_facet | Chang, Heng-Jui Bhati, Saurabhchand Glass, James Liu, Alexander H. |
| contents | Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18843 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | USAD: Universal Speech and Audio Representation via Distillation Chang, Heng-Jui Bhati, Saurabhchand Glass, James Liu, Alexander H. Sound Computation and Language Audio and Speech Processing Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks. |
| title | USAD: Universal Speech and Audio Representation via Distillation |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.18843 |