Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lam, Tsz Kin, Gaido, Marco, Papi, Sara, Bentivogli, Luisa, Haddow, Barry
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910818932621312
author Lam, Tsz Kin
Gaido, Marco
Papi, Sara
Bentivogli, Luisa
Haddow, Barry
author_facet Lam, Tsz Kin
Gaido, Marco
Papi, Sara
Bentivogli, Luisa
Haddow, Barry
contents Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of communication. The most widespread approach to integrating speech into LLMs is dense feature prepending (DFP), which prepends the projected speech representations to the textual representations, allowing end-to-end training with a speech encoder. This raises questions about the need for a sophisticated speech encoder for DFP and how its performance compares with a standard encoder-decoder (i.e., cross-attention) architecture. We compare DFP and cross-attention under a variety of configurations, such as CTC compression, sequence-level knowledge distillation, on monolingual, bilingual, and multilingual models. To perform a controlled architectural comparison, we train all models from scratch rather than using large pretrained models and use comparable data and parameter settings, testing speech-to-text recognition (ASR) and translation (ST) on MuST-C v1.0 and CoVoST2 datasets. Despite the wide adoption of DFP, our results do not indicate a clear advantage of DFP over cross-attention.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02370
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
Lam, Tsz Kin
Gaido, Marco
Papi, Sara
Bentivogli, Luisa
Haddow, Barry
Computation and Language
Sound
Audio and Speech Processing
Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of communication. The most widespread approach to integrating speech into LLMs is dense feature prepending (DFP), which prepends the projected speech representations to the textual representations, allowing end-to-end training with a speech encoder. This raises questions about the need for a sophisticated speech encoder for DFP and how its performance compares with a standard encoder-decoder (i.e., cross-attention) architecture. We compare DFP and cross-attention under a variety of configurations, such as CTC compression, sequence-level knowledge distillation, on monolingual, bilingual, and multilingual models. To perform a controlled architectural comparison, we train all models from scratch rather than using large pretrained models and use comparable data and parameter settings, testing speech-to-text recognition (ASR) and translation (ST) on MuST-C v1.0 and CoVoST2 datasets. Despite the wide adoption of DFP, our results do not indicate a clear advantage of DFP over cross-attention.
title Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.02370