Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Jinheng, Mao, Weijia, Bai, Zechen, Zhang, David Junhao, Wang, Weihao, Lin, Kevin Qinghong, Gu, Yuchao, Chen, Zhijie, Yang, Zhenheng, Shou, Mike Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911141832163328
author Xie, Jinheng
Mao, Weijia
Bai, Zechen
Zhang, David Junhao
Wang, Weihao
Lin, Kevin Qinghong
Gu, Yuchao
Chen, Zhijie
Yang, Zhenheng
Shou, Mike Zheng
author_facet Xie, Jinheng
Mao, Weijia
Bai, Zechen
Zhang, David Junhao
Wang, Weihao
Lin, Kevin Qinghong
Gu, Yuchao
Chen, Zhijie
Yang, Zhenheng
Shou, Mike Zheng
contents We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model. Code and models are released at https://github.com/showlab/Show-o.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12528
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Xie, Jinheng
Mao, Weijia
Bai, Zechen
Zhang, David Junhao
Wang, Weihao
Lin, Kevin Qinghong
Gu, Yuchao
Chen, Zhijie
Yang, Zhenheng
Shou, Mike Zheng
Computer Vision and Pattern Recognition
We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model. Code and models are released at https://github.com/showlab/Show-o.
title Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.12528