Ensuring Consistency for In-Image Translation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fu, Chengpeng, Feng, Xiaocheng, Huang, Yichong, Huo, Wenshuai, Li, Baohang, Zhang, Zhirui, Lu, Yunfei, Tu, Dandan, Tang, Duyu, Wang, Hui, Qin, Bing, Liu, Ting
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915078114115584
author Fu, Chengpeng
Feng, Xiaocheng
Huang, Yichong
Huo, Wenshuai
Li, Baohang
Zhang, Zhirui
Lu, Yunfei
Tu, Dandan
Tang, Duyu
Wang, Hui
Qin, Bing
Liu, Ting
author_facet Fu, Chengpeng
Feng, Xiaocheng
Huang, Yichong
Huo, Wenshuai
Li, Baohang
Zhang, Zhirui
Lu, Yunfei
Tu, Dandan
Tang, Duyu
Wang, Hui
Qin, Bing
Liu, Ting
contents The in-image machine translation task involves translating text embedded within images, with the translated results presented in image format. While this task has numerous applications in various scenarios such as film poster translation and everyday scene image translation, existing methods frequently neglect the aspect of consistency throughout this process. We propose the need to uphold two types of consistency in this task: translation consistency and image generation consistency. The former entails incorporating image information during translation, while the latter involves maintaining consistency between the style of the text-image and the original image, ensuring background integrity. To address these consistency requirements, we introduce a novel two-stage framework named HCIIT (High-Consistency In-Image Translation) which involves text-image translation using a multimodal multilingual large language model in the first stage and image backfilling with a diffusion model in the second stage. Chain of thought learning is utilized in the first stage to enhance the model's ability to leverage image information during translation. Subsequently, a diffusion model trained for style-consistent text-image generation ensures uniformity in text style within images and preserves background details. A dataset comprising 400,000 style-consistent pseudo text-image pairs is curated for model training. Results obtained on both curated test sets and authentic image test sets validate the effectiveness of our framework in ensuring consistency and producing high-quality translated images.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18139
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ensuring Consistency for In-Image Translation
Fu, Chengpeng
Feng, Xiaocheng
Huang, Yichong
Huo, Wenshuai
Li, Baohang
Zhang, Zhirui
Lu, Yunfei
Tu, Dandan
Tang, Duyu
Wang, Hui
Qin, Bing
Liu, Ting
Computation and Language
The in-image machine translation task involves translating text embedded within images, with the translated results presented in image format. While this task has numerous applications in various scenarios such as film poster translation and everyday scene image translation, existing methods frequently neglect the aspect of consistency throughout this process. We propose the need to uphold two types of consistency in this task: translation consistency and image generation consistency. The former entails incorporating image information during translation, while the latter involves maintaining consistency between the style of the text-image and the original image, ensuring background integrity. To address these consistency requirements, we introduce a novel two-stage framework named HCIIT (High-Consistency In-Image Translation) which involves text-image translation using a multimodal multilingual large language model in the first stage and image backfilling with a diffusion model in the second stage. Chain of thought learning is utilized in the first stage to enhance the model's ability to leverage image information during translation. Subsequently, a diffusion model trained for style-consistent text-image generation ensures uniformity in text style within images and preserves background details. A dataset comprising 400,000 style-consistent pseudo text-image pairs is curated for model training. Results obtained on both curated test sets and authentic image test sets validate the effectiveness of our framework in ensuring consistency and producing high-quality translated images.
title Ensuring Consistency for In-Image Translation
topic Computation and Language
url https://arxiv.org/abs/2412.18139