Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shanyuan, Cheng, Bo, Ma, Yuhang, Wu, Liebucha, Ma, Ao, Wu, Xiaoyu, Leng, Dawei, Yin, Yuhui
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917350489456640
author Liu, Shanyuan
Cheng, Bo
Ma, Yuhang
Wu, Liebucha
Ma, Ao
Wu, Xiaoyu
Leng, Dawei
Yin, Yuhui
author_facet Liu, Shanyuan
Cheng, Bo
Ma, Yuhang
Wu, Liebucha
Ma, Ao
Wu, Xiaoyu
Leng, Dawei
Yin, Yuhui
contents Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their English world centric training data, which creates a dilemma for development of other language-native TTI models. One common choice is to fine-tune the English-native TTI model with translated samples. It falls short of fully addressing the model bias problem. Alternatively, training non-English language native models from scratch can effectively resolve the English world bias, but model trained this way would diverge from the English TTI communities, thus not able to utilize the strides continuously gaining in the English TTI communities any more. To build Chinese TTI model meanwhile keep compatibility with the English TTI communities, we propose a novel model structure referred as "Bridge Diffusion Model" (BDM). The proposed BDM employs a backbone-branch network structure to learn the Chinese semantics while keep the latent space compatible with the English-native TTI backbone, in an end-to-end manner. The unique advantages of the proposed BDM are that it's not only adept at generating images that precisely depict Chinese semantics, but also compatible with various English-native TTI plugins, such as different checkpoints, LoRA, ControlNet, Dreambooth, and Textual Inversion, etc. Moreover, BDM can concurrently generate content seamlessly combining both Chinese-native and English-native semantics within a single image, fostering cultural interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2309_00952
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
Liu, Shanyuan
Cheng, Bo
Ma, Yuhang
Wu, Liebucha
Ma, Ao
Wu, Xiaoyu
Leng, Dawei
Yin, Yuhui
Computation and Language
Artificial Intelligence
I.2.10; I.4.9
Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their English world centric training data, which creates a dilemma for development of other language-native TTI models. One common choice is to fine-tune the English-native TTI model with translated samples. It falls short of fully addressing the model bias problem. Alternatively, training non-English language native models from scratch can effectively resolve the English world bias, but model trained this way would diverge from the English TTI communities, thus not able to utilize the strides continuously gaining in the English TTI communities any more. To build Chinese TTI model meanwhile keep compatibility with the English TTI communities, we propose a novel model structure referred as "Bridge Diffusion Model" (BDM). The proposed BDM employs a backbone-branch network structure to learn the Chinese semantics while keep the latent space compatible with the English-native TTI backbone, in an end-to-end manner. The unique advantages of the proposed BDM are that it's not only adept at generating images that precisely depict Chinese semantics, but also compatible with various English-native TTI plugins, such as different checkpoints, LoRA, ControlNet, Dreambooth, and Textual Inversion, etc. Moreover, BDM can concurrently generate content seamlessly combining both Chinese-native and English-native semantics within a single image, fostering cultural interaction.
title Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
topic Computation and Language
Artificial Intelligence
I.2.10; I.4.9
url https://arxiv.org/abs/2309.00952