Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nachmani, Eliya, Levkovitch, Alon, Hirsch, Roy, Salazar, Julian, Asawaroengchai, Chulayuth, Mariooryad, Soroosh, Rivlin, Ehud, Skerry-Ryan, RJ, Ramanovich, Michelle Tadmor
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914816770179072
author Nachmani, Eliya
Levkovitch, Alon
Hirsch, Roy
Salazar, Julian
Asawaroengchai, Chulayuth
Mariooryad, Soroosh
Rivlin, Ehud
Skerry-Ryan, RJ
Ramanovich, Michelle Tadmor
author_facet Nachmani, Eliya
Levkovitch, Alon
Hirsch, Roy
Salazar, Julian
Asawaroengchai, Chulayuth
Mariooryad, Soroosh
Rivlin, Ehud
Skerry-Ryan, RJ
Ramanovich, Michelle Tadmor
contents We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples (https://michelleramanovich.github.io/spectron/spectron) and spoken QA dataset (https://github.com/google-research-datasets/LLAMA1-Test-Set).
format Preprint
id arxiv_https___arxiv_org_abs_2305_15255
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Nachmani, Eliya
Levkovitch, Alon
Hirsch, Roy
Salazar, Julian
Asawaroengchai, Chulayuth
Mariooryad, Soroosh
Rivlin, Ehud
Skerry-Ryan, RJ
Ramanovich, Michelle Tadmor
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples (https://michelleramanovich.github.io/spectron/spectron) and spoken QA dataset (https://github.com/google-research-datasets/LLAMA1-Test-Set).
title Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2305.15255