3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Xinyu, Liu, Xuebo, Wong, Derek F., Rao, Jun, Li, Bei, Ding, Liang, Chao, Lidia S., Tao, Dacheng, Zhang, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911857152884736
author Ma, Xinyu
Liu, Xuebo
Wong, Derek F.
Rao, Jun
Li, Bei
Ding, Liang
Chao, Lidia S.
Tao, Dacheng
Zhang, Min
author_facet Ma, Xinyu
Liu, Xuebo
Wong, Derek F.
Rao, Jun
Li, Bei
Ding, Liang
Chao, Lidia S.
Tao, Dacheng
Zhang, Min
contents Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18413
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
Ma, Xinyu
Liu, Xuebo
Wong, Derek F.
Rao, Jun
Li, Bei
Ding, Liang
Chao, Lidia S.
Tao, Dacheng
Zhang, Min
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
title 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2404.18413