MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Wenyu, Sun, Shuo, Wang, Bin, Zou, Xunlong, Liu, Zhuohan, He, Yingxu, Lin, Geyu, Chen, Nancy F., Aw, Ai Ti
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913800312061952
author Zhang, Wenyu
Sun, Shuo
Wang, Bin
Zou, Xunlong
Liu, Zhuohan
He, Yingxu
Lin, Geyu
Chen, Nancy F.
Aw, Ai Ti
author_facet Zhang, Wenyu
Sun, Shuo
Wang, Bin
Zou, Xunlong
Liu, Zhuohan
He, Yingxu
Lin, Geyu
Chen, Nancy F.
Aw, Ai Ti
contents The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of `weak' encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively light weight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06635
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
Zhang, Wenyu
Sun, Shuo
Wang, Bin
Zou, Xunlong
Liu, Zhuohan
He, Yingxu
Lin, Geyu
Chen, Nancy F.
Aw, Ai Ti
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of `weak' encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively light weight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
title MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2409.06635