Prompting Large Language Models with Audio for General-Purpose Speech Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Wonjune, Roy, Deb
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916391552024576
author Kang, Wonjune
Roy, Deb
author_facet Kang, Wonjune
Roy, Deb
contents In this work, we introduce a framework for speech summarization that leverages the processing and reasoning capabilities of large language models (LLMs). We propose an end-to-end system that combines an instruction-tuned LLM with an audio encoder that converts speech into token representations that the LLM can interpret. Using a dataset with paired speech-text data, the overall system is trained to generate consistent responses to prompts with the same semantic information regardless of the input modality. The resulting framework allows the LLM to process speech inputs in the same way as text, enabling speech summarization by simply prompting the LLM. Unlike prior approaches, our method is able to summarize spoken content from any arbitrary domain, and it can produce summaries in different styles by varying the LLM prompting strategy. Experiments demonstrate that our approach outperforms a cascade baseline of speech recognition followed by LLM text processing.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05968
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompting Large Language Models with Audio for General-Purpose Speech Summarization
Kang, Wonjune
Roy, Deb
Audio and Speech Processing
Computation and Language
In this work, we introduce a framework for speech summarization that leverages the processing and reasoning capabilities of large language models (LLMs). We propose an end-to-end system that combines an instruction-tuned LLM with an audio encoder that converts speech into token representations that the LLM can interpret. Using a dataset with paired speech-text data, the overall system is trained to generate consistent responses to prompts with the same semantic information regardless of the input modality. The resulting framework allows the LLM to process speech inputs in the same way as text, enabling speech summarization by simply prompting the LLM. Unlike prior approaches, our method is able to summarize spoken content from any arbitrary domain, and it can produce summaries in different styles by varying the LLM prompting strategy. Experiments demonstrate that our approach outperforms a cascade baseline of speech recognition followed by LLM text processing.
title Prompting Large Language Models with Audio for General-Purpose Speech Summarization
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2406.05968