Saved in:
Bibliographic Details
Main Authors: Vijayan, Vipin, Bowen, Braeden, Grigsby, Scott, Anderson, Timothy, Gwinnup, Jeremy
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.03045
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916147517980672
author Vijayan, Vipin
Bowen, Braeden
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
author_facet Vijayan, Vipin
Bowen, Braeden
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
contents While most current work in multimodal machine translation (MMT) uses the Multi30k dataset for training and evaluation, we find that the resulting models overfit to the Multi30k dataset to an extreme degree. Consequently, these models perform very badly when evaluated against typical text-only testing sets such as the WMT newstest datasets. In order to perform well on both Multi30k and typical text-only datasets, we use a performant text-only machine translation (MT) model as the starting point of our MMT model. We add vision-text adapter layers connected via gating mechanisms to the MT model, and incrementally transform the MT model into an MMT model by 1) pre-training using vision-based masking of the source text and 2) fine-tuning on Multi30k.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03045
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adding Multimodal Capabilities to a Text-only Translation Model
Vijayan, Vipin
Bowen, Braeden
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
Computation and Language
While most current work in multimodal machine translation (MMT) uses the Multi30k dataset for training and evaluation, we find that the resulting models overfit to the Multi30k dataset to an extreme degree. Consequently, these models perform very badly when evaluated against typical text-only testing sets such as the WMT newstest datasets. In order to perform well on both Multi30k and typical text-only datasets, we use a performant text-only machine translation (MT) model as the starting point of our MMT model. We add vision-text adapter layers connected via gating mechanisms to the MT model, and incrementally transform the MT model into an MMT model by 1) pre-training using vision-based masking of the source text and 2) fine-tuning on Multi30k.
title Adding Multimodal Capabilities to a Text-only Translation Model
topic Computation and Language
url https://arxiv.org/abs/2403.03045