STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shen, Ying, Chen, Tianrong, Gao, Yuan, Zhang, Yizhe, Wang, Yuyang, Bautista, Miguel Ángel, Zhai, Shuangfei, Susskind, Joshua M., Gu, Jiatao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914545388224512
author Shen, Ying
Chen, Tianrong
Gao, Yuan
Zhang, Yizhe
Wang, Yuyang
Bautista, Miguel Ángel
Zhai, Shuangfei
Susskind, Joshua M.
Gu, Jiatao
author_facet Shen, Ying
Chen, Tianrong
Gao, Yuan
Zhang, Yizhe
Wang, Yuyang
Bautista, Miguel Ángel
Zhai, Shuangfei
Susskind, Joshua M.
Gu, Jiatao
contents Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers--sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs--making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08029
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Shen, Ying
Chen, Tianrong
Gao, Yuan
Zhang, Yizhe
Wang, Yuyang
Bautista, Miguel Ángel
Zhai, Shuangfei
Susskind, Joshua M.
Gu, Jiatao
Computer Vision and Pattern Recognition
Machine Learning
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers--sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs--making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.
title STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.08029