Saved in:
Bibliographic Details
Main Authors: Bowen, Braeden, Vijayan, Vipin, Grigsby, Scott, Anderson, Timothy, Gwinnup, Jeremy
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.03075
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911789545947136
author Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
author_facet Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
contents The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce new methods for detection of visually and contextually relevant (concrete) tokens from source sentences, including detection with natural language processing (NLP), detection with object detection, and a joint detection-verification technique. We also introduce new methods for selection of detected tokens, including shortest $n$ tokens, longest $n$ tokens, and all detected concrete tokens. We utilize the GRAM MMT architecture to train models against synthetically collated multimodal datasets of source images with masked sentences, showing performance improvements and improved usage of visual context during translation tasks over the baseline model.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03075
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detecting Concrete Visual Tokens for Multimodal Machine Translation
Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
Computation and Language
The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce new methods for detection of visually and contextually relevant (concrete) tokens from source sentences, including detection with natural language processing (NLP), detection with object detection, and a joint detection-verification technique. We also introduce new methods for selection of detected tokens, including shortest $n$ tokens, longest $n$ tokens, and all detected concrete tokens. We utilize the GRAM MMT architecture to train models against synthetically collated multimodal datasets of source images with masked sentences, showing performance improvements and improved usage of visual context during translation tasks over the baseline model.
title Detecting Concrete Visual Tokens for Multimodal Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2403.03075