Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Andong, Song, Yuchen, Chen, Kehai, Yang, Muyun, Zhao, Tiejun, Zhang, Min
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909448429109248
author Chen, Andong
Song, Yuchen
Chen, Kehai
Yang, Muyun
Zhao, Tiejun
Zhang, Min
author_facet Chen, Andong
Song, Yuchen
Chen, Kehai
Yang, Muyun
Zhao, Tiejun
Zhang, Min
contents Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we introduce a stable diffusion-based imagination network into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing the multimodel MT. Particularly, we build heuristic human feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of image annotation, which breaks the bottleneck of using visual information in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into large-scale text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 14 BLEU points on Multi30K multimodal MT benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12627
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
Chen, Andong
Song, Yuchen
Chen, Kehai
Yang, Muyun
Zhao, Tiejun
Zhang, Min
Computation and Language
Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we introduce a stable diffusion-based imagination network into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing the multimodel MT. Particularly, we build heuristic human feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of image annotation, which breaks the bottleneck of using visual information in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into large-scale text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 14 BLEU points on Multi30K multimodal MT benchmarks.
title Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2412.12627