Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.03045 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916147517980672 |
|---|---|
| author | Vijayan, Vipin Bowen, Braeden Grigsby, Scott Anderson, Timothy Gwinnup, Jeremy |
| author_facet | Vijayan, Vipin Bowen, Braeden Grigsby, Scott Anderson, Timothy Gwinnup, Jeremy |
| contents | While most current work in multimodal machine translation (MMT) uses the Multi30k dataset for training and evaluation, we find that the resulting models overfit to the Multi30k dataset to an extreme degree. Consequently, these models perform very badly when evaluated against typical text-only testing sets such as the WMT newstest datasets. In order to perform well on both Multi30k and typical text-only datasets, we use a performant text-only machine translation (MT) model as the starting point of our MMT model. We add vision-text adapter layers connected via gating mechanisms to the MT model, and incrementally transform the MT model into an MMT model by 1) pre-training using vision-based masking of the source text and 2) fine-tuning on Multi30k. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_03045 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Adding Multimodal Capabilities to a Text-only Translation Model Vijayan, Vipin Bowen, Braeden Grigsby, Scott Anderson, Timothy Gwinnup, Jeremy Computation and Language While most current work in multimodal machine translation (MMT) uses the Multi30k dataset for training and evaluation, we find that the resulting models overfit to the Multi30k dataset to an extreme degree. Consequently, these models perform very badly when evaluated against typical text-only testing sets such as the WMT newstest datasets. In order to perform well on both Multi30k and typical text-only datasets, we use a performant text-only machine translation (MT) model as the starting point of our MMT model. We add vision-text adapter layers connected via gating mechanisms to the MT model, and incrementally transform the MT model into an MMT model by 1) pre-training using vision-based masking of the source text and 2) fine-tuning on Multi30k. |
| title | Adding Multimodal Capabilities to a Text-only Translation Model |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2403.03045 |