Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Xichen, Dong, Li, Huang, Shaohan, Peng, Zhiliang, Chen, Wenhu, Wei, Furu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917650833080320
author Pan, Xichen
Dong, Li
Huang, Shaohan
Peng, Zhiliang
Chen, Wenhu
Wei, Furu
author_facet Pan, Xichen
Dong, Li
Huang, Shaohan
Peng, Zhiliang
Chen, Wenhu
Wei, Furu
contents Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." The code can be found at https://aka.ms/Kosmos-G
format Preprint
id arxiv_https___arxiv_org_abs_2310_02992
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Kosmos-G: Generating Images in Context with Multimodal Large Language Models
Pan, Xichen
Dong, Li
Huang, Shaohan
Peng, Zhiliang
Chen, Wenhu
Wei, Furu
Computer Vision and Pattern Recognition
Computation and Language
Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." The code can be found at https://aka.ms/Kosmos-G
title Kosmos-G: Generating Images in Context with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2310.02992