Saved in:
Bibliographic Details
Main Authors: Ma, Yiyang, Liu, Xingchao, Chen, Xiaokang, Liu, Wen, Wu, Chengyue, Wu, Zhiyu, Pan, Zizheng, Xie, Zhenda, Zhang, Haowei, yu, Xingkai, Zhao, Liang, Wang, Yisong, Liu, Jiaying, Ruan, Chong
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.07975
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917965821116416
author Ma, Yiyang
Liu, Xingchao
Chen, Xiaokang
Liu, Wen
Wu, Chengyue
Wu, Zhiyu
Pan, Zizheng
Xie, Zhenda
Zhang, Haowei
yu, Xingkai
Zhao, Liang
Wang, Yisong
Liu, Jiaying
Ruan, Chong
author_facet Ma, Yiyang
Liu, Xingchao
Chen, Xiaokang
Liu, Wen
Wu, Chengyue
Wu, Zhiyu
Pan, Zizheng
Xie, Zhenda
Zhang, Haowei
yu, Xingkai
Zhao, Liang
Wang, Yisong
Liu, Jiaying
Ruan, Chong
contents We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches across standard benchmarks. This work represents a step toward more efficient and versatile vision-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07975
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Ma, Yiyang
Liu, Xingchao
Chen, Xiaokang
Liu, Wen
Wu, Chengyue
Wu, Zhiyu
Pan, Zizheng
Xie, Zhenda
Zhang, Haowei
yu, Xingkai
Zhao, Liang
Wang, Yisong
Liu, Jiaying
Ruan, Chong
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches across standard benchmarks. This work represents a step toward more efficient and versatile vision-language models.
title JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.07975