Distilling a speech and music encoder with task arithmetic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ritter-Gutierrez, Fabian, Lin, Yi-Cheng, Wei, Jui-Chiang, Wong, Jeremy H. M, Chng, Eng Siong, Chen, Nancy F., Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909615883550720
author Ritter-Gutierrez, Fabian
Lin, Yi-Cheng
Wei, Jui-Chiang
Wong, Jeremy H. M
Chng, Eng Siong
Chen, Nancy F.
Lee, Hung-yi
author_facet Ritter-Gutierrez, Fabian
Lin, Yi-Cheng
Wei, Jui-Chiang
Wong, Jeremy H. M
Chng, Eng Siong
Chen, Nancy F.
Lee, Hung-yi
contents Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distilling a speech and music encoder with task arithmetic
Ritter-Gutierrez, Fabian
Lin, Yi-Cheng
Wei, Jui-Chiang
Wong, Jeremy H. M
Chng, Eng Siong
Chen, Nancy F.
Lee, Hung-yi
Sound
Audio and Speech Processing
Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.
title Distilling a speech and music encoder with task arithmetic
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.13270