Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jinyi, Yao, Yuan, Wang, Chongyi, Wang, Shan, Pan, Yinxu, Chen, Qianyu, Yu, Tianyu, Wu, Hanghao, Zhao, Yue, Zhang, Haoye, Han, Xu, Lin, Yankai, Xue, Jiao, Li, Dahai, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929285638389760
author Hu, Jinyi
Yao, Yuan
Wang, Chongyi
Wang, Shan
Pan, Yinxu
Chen, Qianyu
Yu, Tianyu
Wu, Hanghao
Zhao, Yue
Zhang, Haoye
Han, Xu
Lin, Yankai
Xue, Jiao
Li, Dahai
Liu, Zhiyuan
Sun, Maosong
author_facet Hu, Jinyi
Yao, Yuan
Wang, Chongyi
Wang, Shan
Pan, Yinxu
Chen, Qianyu
Yu, Tianyu
Wu, Hanghao
Zhao, Yue
Zhang, Haoye
Han, Xu
Lin, Yankai
Xue, Jiao
Li, Dahai
Liu, Zhiyuan
Sun, Maosong
contents Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.e., lack of large-scale, high-quality image-text data). In this work, we propose MPM, an effective training paradigm for training large multimodal models in non-English languages. MPM demonstrates that Multilingual language models can Pivot zero-shot Multimodal learning across languages. Specifically, based on a strong multilingual large language model, multimodal models pretrained on English-only image-text data can well generalize to other languages in a (quasi)-zero-shot manner, even surpassing models trained on image-text data in native languages. Taking Chinese as a practice of MPM, we build large multimodal models VisCPM in image-to-text and text-to-image generation, which achieve state-of-the-art (open-source) performance in Chinese. To facilitate future research, we open-source codes and model weights at https://github.com/OpenBMB/VisCPM.git.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12038
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
Hu, Jinyi
Yao, Yuan
Wang, Chongyi
Wang, Shan
Pan, Yinxu
Chen, Qianyu
Yu, Tianyu
Wu, Hanghao
Zhao, Yue
Zhang, Haoye
Han, Xu
Lin, Yankai
Xue, Jiao
Li, Dahai
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Computer Vision and Pattern Recognition
Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.e., lack of large-scale, high-quality image-text data). In this work, we propose MPM, an effective training paradigm for training large multimodal models in non-English languages. MPM demonstrates that Multilingual language models can Pivot zero-shot Multimodal learning across languages. Specifically, based on a strong multilingual large language model, multimodal models pretrained on English-only image-text data can well generalize to other languages in a (quasi)-zero-shot manner, even surpassing models trained on image-text data in native languages. Taking Chinese as a practice of MPM, we build large multimodal models VisCPM in image-to-text and text-to-image generation, which achieve state-of-the-art (open-source) performance in Chinese. To facilitate future research, we open-source codes and model weights at https://github.com/OpenBMB/VisCPM.git.
title Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.12038