Aligning Pre-trained Models for Spoken Language Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sedláček, Šimon, Kesiraju, Santosh, Polok, Alexander, Černocký, Jan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915037244817408
author Sedláček, Šimon
Kesiraju, Santosh
Polok, Alexander
Černocký, Jan
author_facet Sedláček, Šimon
Kesiraju, Santosh
Polok, Alexander
Černocký, Jan
contents This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT) models via a small connector module (Q-Former, our Subsampler-Transformer Encoder). This connector bridges the gap between the speech and text modalities, transforming ASR encoder embeddings into the latent representation space of the MT encoder while being the only part of the system optimized during training. Experiments are conducted on the How2 English-Portuguese dataset as we investigate the alignment approach in a small-scale scenario focusing on ST. While keeping the size of the connector module constant and small in comparison ( < 5% of the size of the larger aligned models), increasing the size and capability of the foundation ASR and MT models universally improves translation results. We also find that the connectors can serve as domain adapters for the foundation MT models, significantly improving translation performance in the aligned ST setting. We conclude that this approach represents a viable and scalable approach to training end-to-end ST systems.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18294
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aligning Pre-trained Models for Spoken Language Translation
Sedláček, Šimon
Kesiraju, Santosh
Polok, Alexander
Černocký, Jan
Computation and Language
Artificial Intelligence
Machine Learning
This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT) models via a small connector module (Q-Former, our Subsampler-Transformer Encoder). This connector bridges the gap between the speech and text modalities, transforming ASR encoder embeddings into the latent representation space of the MT encoder while being the only part of the system optimized during training. Experiments are conducted on the How2 English-Portuguese dataset as we investigate the alignment approach in a small-scale scenario focusing on ST. While keeping the size of the connector module constant and small in comparison ( < 5% of the size of the larger aligned models), increasing the size and capability of the foundation ASR and MT models universally improves translation results. We also find that the connectors can serve as domain adapters for the foundation MT models, significantly improving translation performance in the aligned ST setting. We conclude that this approach represents a viable and scalable approach to training end-to-end ST systems.
title Aligning Pre-trained Models for Spoken Language Translation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.18294