Detecting Concrete Visual Tokens for Multimodal Machine Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bowen, Braeden, Vijayan, Vipin, Grigsby, Scott, Anderson, Timothy, Gwinnup, Jeremy
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911789545947136
author Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
author_facet Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
contents The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce new methods for detection of visually and contextually relevant (concrete) tokens from source sentences, including detection with natural language processing (NLP), detection with object detection, and a joint detection-verification technique. We also introduce new methods for selection of detected tokens, including shortest $n$ tokens, longest $n$ tokens, and all detected concrete tokens. We utilize the GRAM MMT architecture to train models against synthetically collated multimodal datasets of source images with masked sentences, showing performance improvements and improved usage of visual context during translation tasks over the baseline model.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03075
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detecting Concrete Visual Tokens for Multimodal Machine Translation
Bowen, Braeden
Vijayan, Vipin
Grigsby, Scott
Anderson, Timothy
Gwinnup, Jeremy
Computation and Language
The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce new methods for detection of visually and contextually relevant (concrete) tokens from source sentences, including detection with natural language processing (NLP), detection with object detection, and a joint detection-verification technique. We also introduce new methods for selection of detected tokens, including shortest $n$ tokens, longest $n$ tokens, and all detected concrete tokens. We utilize the GRAM MMT architecture to train models against synthetically collated multimodal datasets of source images with masked sentences, showing performance improvements and improved usage of visual context during translation tasks over the baseline model.
title Detecting Concrete Visual Tokens for Multimodal Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2403.03075