Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Bingchen, Guo, Qiushan, Wang, Ye, Huang, Yixuan, Zhai, Zhonghua, Tian, Yu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917244524560384
author Zhao, Bingchen
Guo, Qiushan
Wang, Ye
Huang, Yixuan
Zhai, Zhonghua
Tian, Yu
author_facet Zhao, Bingchen
Guo, Qiushan
Wang, Ye
Huang, Yixuan
Zhai, Zhonghua
Tian, Yu
contents We introduce CompTok, a training framework for learning visual tokenizers whose tokens are enhanced for compositionality. CompTok uses a token-conditioned diffusion decoder. By employing an InfoGAN-style objective, where we train a recognition model to predict the tokens used to condition the diffusion decoder using the decoded images, we enforce the decoder to not ignore any of the tokens. To promote compositional control, besides the original images, CompTok also trains on tokens formed by swapping token subsets between images, enabling more compositional control of the token over the decoder. As the swapped tokens between images do not have ground truth image targets, we apply a manifold constraint via an adversarial flow regularizer to keep unpaired swap generations on the natural-image distribution. The resulting tokenizer not only achieves state-of-the-art performance on image class-conditioned generation, but also demonstrates properties such as swapping tokens between images to achieve high level semantic editing of an image. Additionally, we propose two metrics that measures the landscape of the token space that can be useful to describe not only the compositionality of the tokens, but also how easy to learn the landscape is for a generator to be trained on this space. We show in experiments that CompTok can improve on both of the metrics as well as supporting state-of-the-art generators for class conditioned generation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03339
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
Zhao, Bingchen
Guo, Qiushan
Wang, Ye
Huang, Yixuan
Zhai, Zhonghua
Tian, Yu
Computer Vision and Pattern Recognition
We introduce CompTok, a training framework for learning visual tokenizers whose tokens are enhanced for compositionality. CompTok uses a token-conditioned diffusion decoder. By employing an InfoGAN-style objective, where we train a recognition model to predict the tokens used to condition the diffusion decoder using the decoded images, we enforce the decoder to not ignore any of the tokens. To promote compositional control, besides the original images, CompTok also trains on tokens formed by swapping token subsets between images, enabling more compositional control of the token over the decoder. As the swapped tokens between images do not have ground truth image targets, we apply a manifold constraint via an adversarial flow regularizer to keep unpaired swap generations on the natural-image distribution. The resulting tokenizer not only achieves state-of-the-art performance on image class-conditioned generation, but also demonstrates properties such as swapping tokens between images to achieve high level semantic editing of an image. Additionally, we propose two metrics that measures the landscape of the token space that can be useful to describe not only the compositionality of the tokens, but also how easy to learn the landscape is for a generator to be trained on this space. We show in experiments that CompTok can improve on both of the metrics as well as supporting state-of-the-art generators for class conditioned generation.
title Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03339