Transferable speech-to-text large language model alignment module

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Boyong, Yan, Chao, Pu, Haoran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916295282262016
author Wu, Boyong
Yan, Chao
Pu, Haoran
author_facet Wu, Boyong
Yan, Chao
Pu, Haoran
contents By leveraging the power of Large Language Models(LLMs) and speech foundation models, state of the art speech-text bimodal works can achieve challenging tasks like spoken translation(ST) and question answering(SQA) altogether with much simpler architectures. In this paper, we utilize the capability of Whisper encoder and pre-trained Yi-6B. Empirical results reveal that modal alignment can be achieved with one layer module and hundred hours of speech-text multitask corpus. We further swap the Yi-6B with human preferences aligned version of Yi-6B-Chat during inference, and discover that the alignment capability is applicable as well. In addition, the alignment subspace revealed by singular value decomposition(SVD) also implies linear alignment subspace is sparse, which leaves the possibility to concatenate other features like voice-print or video to expand modality.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13357
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transferable speech-to-text large language model alignment module
Wu, Boyong
Yan, Chao
Pu, Haoran
Computation and Language
Sound
Audio and Speech Processing
By leveraging the power of Large Language Models(LLMs) and speech foundation models, state of the art speech-text bimodal works can achieve challenging tasks like spoken translation(ST) and question answering(SQA) altogether with much simpler architectures. In this paper, we utilize the capability of Whisper encoder and pre-trained Yi-6B. Empirical results reveal that modal alignment can be achieved with one layer module and hundred hours of speech-text multitask corpus. We further swap the Yi-6B with human preferences aligned version of Yi-6B-Chat during inference, and discover that the alignment capability is applicable as well. In addition, the alignment subspace revealed by singular value decomposition(SVD) also implies linear alignment subspace is sparse, which leaves the possibility to concatenate other features like voice-print or video to expand modality.
title Transferable speech-to-text large language model alignment module
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.13357