Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chengyue, Chen, Xiaokang, Wu, Zhiyu, Ma, Yiyang, Liu, Xingchao, Pan, Zizheng, Liu, Wen, Xie, Zhenda, Yu, Xingkai, Ruan, Chong, Luo, Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929548169314304
author Wu, Chengyue
Chen, Xiaokang
Wu, Zhiyu
Ma, Yiyang
Liu, Xingchao
Pan, Zizheng
Liu, Wen
Xie, Zhenda
Yu, Xingkai
Ruan, Chong
Luo, Ping
author_facet Wu, Chengyue
Chen, Xiaokang
Wu, Zhiyu
Ma, Yiyang
Liu, Xingchao
Pan, Zizheng
Liu, Wen
Xie, Zhenda
Yu, Xingkai
Ruan, Chong
Luo, Ping
contents In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Wu, Chengyue
Chen, Xiaokang
Wu, Zhiyu
Ma, Yiyang
Liu, Xingchao
Pan, Zizheng
Liu, Wen
Xie, Zhenda
Yu, Xingkai
Ruan, Chong
Luo, Ping
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.
title Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.13848