CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yitong, Wu, Zuxuan, Qiu, Xipeng, Jiang, Yu-Gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910043835727872
author Chen, Yitong
Wu, Zuxuan
Qiu, Xipeng
Jiang, Yu-Gang
author_facet Chen, Yitong
Wu, Zuxuan
Qiu, Xipeng
Jiang, Yu-Gang
contents Autoregressive (AR) language models rely on causal tokenization, but extending this paradigm to vision remains non-trivial. Current visual tokenizers either flatten 2D patches into non-causal sequences or enforce heuristic orderings that misalign with the "next-token prediction" pattern. Recent diffusion autoencoders similarly fall short: conditioning the decoder on all tokens lacks causality, while applying nested dropout mechanism introduces imbalance. To address these challenges, we present CaTok, a 1D causal image tokenizer with a MeanFlow decoder. By selecting tokens over time intervals and binding them to the MeanFlow objective, as illustrated in Fig. 1, CaTok learns causal 1D representations that support both fast one-step generation and high-fidelity multi-step sampling, while naturally capturing diverse visual concepts across token intervals. To further stabilize and accelerate training, we propose a straightforward regularization REPA-A, which aligns encoder features with Vision Foundation Models (VFMs). Experiments demonstrate that CaTok achieves state-of-the-art results on ImageNet reconstruction, reaching 0.75 FID, 22.53 PSNR and 0.674 SSIM with fewer training epochs, and the AR model attains performance comparable to leading approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06449
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
Chen, Yitong
Wu, Zuxuan
Qiu, Xipeng
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Autoregressive (AR) language models rely on causal tokenization, but extending this paradigm to vision remains non-trivial. Current visual tokenizers either flatten 2D patches into non-causal sequences or enforce heuristic orderings that misalign with the "next-token prediction" pattern. Recent diffusion autoencoders similarly fall short: conditioning the decoder on all tokens lacks causality, while applying nested dropout mechanism introduces imbalance. To address these challenges, we present CaTok, a 1D causal image tokenizer with a MeanFlow decoder. By selecting tokens over time intervals and binding them to the MeanFlow objective, as illustrated in Fig. 1, CaTok learns causal 1D representations that support both fast one-step generation and high-fidelity multi-step sampling, while naturally capturing diverse visual concepts across token intervals. To further stabilize and accelerate training, we propose a straightforward regularization REPA-A, which aligns encoder features with Vision Foundation Models (VFMs). Experiments demonstrate that CaTok achieves state-of-the-art results on ImageNet reconstruction, reaching 0.75 FID, 22.53 PSNR and 0.674 SSIM with fewer training epochs, and the AR model attains performance comparable to leading approaches.
title CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.06449