USAD: Universal Speech and Audio Representation via Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Heng-Jui, Bhati, Saurabhchand, Glass, James, Liu, Alexander H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908492584976384
author Chang, Heng-Jui
Bhati, Saurabhchand
Glass, James
Liu, Alexander H.
author_facet Chang, Heng-Jui
Bhati, Saurabhchand
Glass, James
Liu, Alexander H.
contents Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle USAD: Universal Speech and Audio Representation via Distillation
Chang, Heng-Jui
Bhati, Saurabhchand
Glass, James
Liu, Alexander H.
Sound
Computation and Language
Audio and Speech Processing
Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.
title USAD: Universal Speech and Audio Representation via Distillation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.18843