Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Buettner, Kyle, Kovashka, Adriana
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914967817551872
author Buettner, Kyle
Kovashka, Adriana
author_facet Buettner, Kyle
Kovashka, Adriana
contents There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, multilingual retrieval case study, we quantify the existing lack of model flexibility. We empirically show performance gaps between training on captions that come from native German perception and captions that have been either machine-translated or human-translated from English into German. To address these gaps, we further propose and evaluate caption augmentation strategies. While we achieve mean recall improvements (+1.3), gaps still remain, indicating an open area of future work for the community.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
Buettner, Kyle
Kovashka, Adriana
Computer Vision and Pattern Recognition
Artificial Intelligence
There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, multilingual retrieval case study, we quantify the existing lack of model flexibility. We empirically show performance gaps between training on captions that come from native German perception and captions that have been either machine-translated or human-translated from English into German. To address these gaps, we further propose and evaluate caption augmentation strategies. While we achieve mean recall improvements (+1.3), gaps still remain, indicating an open area of future work for the community.
title Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.02027