Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kou, Siqi, Jin, Jiachun, Liu, Zhihong, Liu, Chang, Ma, Ye, Jia, Jian, Chen, Quan, Jiang, Peng, Deng, Zhijie
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915244641615872
author Kou, Siqi
Jin, Jiachun
Liu, Zhihong
Liu, Chang
Ma, Ye
Jia, Jian
Chen, Quan
Jiang, Peng
Deng, Zhijie
author_facet Kou, Siqi
Jin, Jiachun
Liu, Zhihong
Liu, Chang
Ma, Ye
Jia, Jian
Chen, Quan
Jiang, Peng
Deng, Zhijie
contents We introduce Orthus, an autoregressive (AR) transformer that excels in generating images given textual prompts, answering questions based on visual inputs, and even crafting lengthy image-text interleaved contents. Unlike prior arts on unified multimodal modeling, Orthus simultaneously copes with discrete text tokens and continuous image features under the AR modeling principle. The continuous treatment of visual signals minimizes the information loss for both image understanding and generation while the fully AR formulation renders the characterization of the correlation between modalities straightforward. The key mechanism enabling Orthus to leverage these advantages lies in its modality-specific heads -- one regular language modeling (LM) head predicts discrete text tokens and one diffusion head generates continuous image features conditioning on the output of the backbone. We devise an efficient strategy for building Orthus -- by substituting the Vector Quantization (VQ) operation in the existing unified AR model with a soft alternative, introducing a diffusion head, and tuning the added modules to reconstruct images, we can create an Orthus-base model effortlessly (e.g., within mere 72 A100 GPU hours). Orthus-base can further embrace post-training to better model interleaved images and texts. Empirically, Orthus surpasses competing baselines including Show-o and Chameleon across standard benchmarks, achieving a GenEval score of 0.58 and an MME-P score of 1265.8 using 7B parameters. Orthus also shows exceptional mixed-modality generation capabilities, reflecting the potential for handling intricate practical generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Kou, Siqi
Jin, Jiachun
Liu, Zhihong
Liu, Chang
Ma, Ye
Jia, Jian
Chen, Quan
Jiang, Peng
Deng, Zhijie
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
We introduce Orthus, an autoregressive (AR) transformer that excels in generating images given textual prompts, answering questions based on visual inputs, and even crafting lengthy image-text interleaved contents. Unlike prior arts on unified multimodal modeling, Orthus simultaneously copes with discrete text tokens and continuous image features under the AR modeling principle. The continuous treatment of visual signals minimizes the information loss for both image understanding and generation while the fully AR formulation renders the characterization of the correlation between modalities straightforward. The key mechanism enabling Orthus to leverage these advantages lies in its modality-specific heads -- one regular language modeling (LM) head predicts discrete text tokens and one diffusion head generates continuous image features conditioning on the output of the backbone. We devise an efficient strategy for building Orthus -- by substituting the Vector Quantization (VQ) operation in the existing unified AR model with a soft alternative, introducing a diffusion head, and tuning the added modules to reconstruct images, we can create an Orthus-base model effortlessly (e.g., within mere 72 A100 GPU hours). Orthus-base can further embrace post-training to better model interleaved images and texts. Empirically, Orthus surpasses competing baselines including Show-o and Chameleon across standard benchmarks, achieving a GenEval score of 0.58 and an MME-P score of 1265.8 using 7B parameters. Orthus also shows exceptional mixed-modality generation capabilities, reflecting the potential for handling intricate practical generation tasks.
title Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.00127