UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Zhengrong, Zhang, Haiyu, Zeng, Xiangyu, Chen, Boyu, Wang, Chenting, Zhuang, Shaobin, Dong, Lu, Wang, Yi, Wang, Limin, Wang, Yali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915823997681664
author Yue, Zhengrong
Zhang, Haiyu
Zeng, Xiangyu
Chen, Boyu
Wang, Chenting
Zhuang, Shaobin
Dong, Lu
Wang, Yi
Wang, Limin
Wang, Yali
author_facet Yue, Zhengrong
Zhang, Haiyu
Zeng, Xiangyu
Chen, Boyu
Wang, Chenting
Zhuang, Shaobin
Dong, Lu
Wang, Yi
Wang, Limin
Wang, Yali
contents Tokenizer is a crucial component for both visual understanding and generation. To advance toward the ultimate goal of universal modeling, recent research has focused on developing a unified tokenizer. However, existing tokenizers face a significant performance trade-off between understanding and generation, stemming from the inherent conflict between high-level semantic abstraction and low-level pixel reconstruction. To tackle this challenge, we propose a generic and unified tokenizer, namely UniFlow, by flexibly adapting any visual encoder with a concise reconstruction decoder. Specifically, we introduce layer-wise adaptive self-distillation applied to the well-pretrained visual encoders, which enables UniFlow to simultaneously inherit the strong semantic features for visual understanding and flexibly adapt to model fine-grained details for visual generation. Moreover, we propose a lightweight patch-wise pixel flow decoder, which efficiently achieves high-fidelity pixel reconstruction by modeling a conditional flow from the noisy state back to the patch-wise pixel domain. By leveraging the semantic features as visual conditions for the decoder, we effectively alleviate the training conflicts between understanding and generation. Furthermore, the patch-wise learning strategy simplifies the data distribution, thereby improving training efficiency. Extensive experiments across 13 challenging benchmarks spanning 7 widely studied visual understanding and generation tasks demonstrate that UniFlow achieves a win-win outcome. For instance, our 7B UniFlow-XL not only surpasses the 14B TokenFlow-XL by 6.05% on average understanding benchmarks, but also achieves a competitive results in both visual reconstruction and generation, surpassing UniTok by 0.15 in rFID and 0.09 in gFID (without guidance), respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
Yue, Zhengrong
Zhang, Haiyu
Zeng, Xiangyu
Chen, Boyu
Wang, Chenting
Zhuang, Shaobin
Dong, Lu
Wang, Yi
Wang, Limin
Wang, Yali
Computer Vision and Pattern Recognition
Tokenizer is a crucial component for both visual understanding and generation. To advance toward the ultimate goal of universal modeling, recent research has focused on developing a unified tokenizer. However, existing tokenizers face a significant performance trade-off between understanding and generation, stemming from the inherent conflict between high-level semantic abstraction and low-level pixel reconstruction. To tackle this challenge, we propose a generic and unified tokenizer, namely UniFlow, by flexibly adapting any visual encoder with a concise reconstruction decoder. Specifically, we introduce layer-wise adaptive self-distillation applied to the well-pretrained visual encoders, which enables UniFlow to simultaneously inherit the strong semantic features for visual understanding and flexibly adapt to model fine-grained details for visual generation. Moreover, we propose a lightweight patch-wise pixel flow decoder, which efficiently achieves high-fidelity pixel reconstruction by modeling a conditional flow from the noisy state back to the patch-wise pixel domain. By leveraging the semantic features as visual conditions for the decoder, we effectively alleviate the training conflicts between understanding and generation. Furthermore, the patch-wise learning strategy simplifies the data distribution, thereby improving training efficiency. Extensive experiments across 13 challenging benchmarks spanning 7 widely studied visual understanding and generation tasks demonstrate that UniFlow achieves a win-win outcome. For instance, our 7B UniFlow-XL not only surpasses the 14B TokenFlow-XL by 6.05% on average understanding benchmarks, but also achieves a competitive results in both visual reconstruction and generation, surpassing UniTok by 0.15 in rFID and 0.09 in gFID (without guidance), respectively.
title UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10575