Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Bohan, Yue, Zhongqi, Zhang, Fengda, Chen, Shuo, Bi, Li'an, Zhang, Junzhe, Song, Xue, Chan, Kennard Yanting, Pan, Jiachun, Wu, Weijia, Zhou, Mingze, Lin, Wang, Pan, Kaihang, Zhang, Saining, Jia, Liyu, Hu, Wentao, Zhao, Wei, Zhang, Hanwang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2505.07538
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913860837965824
author Wang, Bohan
Yue, Zhongqi
Zhang, Fengda
Chen, Shuo
Bi, Li'an
Zhang, Junzhe
Song, Xue
Chan, Kennard Yanting
Pan, Jiachun
Wu, Weijia
Zhou, Mingze
Lin, Wang
Pan, Kaihang
Zhang, Saining
Jia, Liyu
Hu, Wentao
Zhao, Wei
Zhang, Hanwang
author_facet Wang, Bohan
Yue, Zhongqi
Zhang, Fengda
Chen, Shuo
Bi, Li'an
Zhang, Junzhe
Song, Xue
Chan, Kennard Yanting
Pan, Jiachun
Wu, Weijia
Zhou, Mingze
Lin, Wang
Pan, Kaihang
Zhang, Saining
Jia, Liyu
Hu, Wentao
Zhao, Wei
Zhang, Hanwang
contents We completely discard the conventional spatial prior in image representation and introduce a novel discrete visual tokenizer: Self-consistency Tokenizer (Selftok). At its design core, we compose an autoregressive (AR) prior -- mirroring the causal structure of language -- into visual tokens by using the reverse diffusion process of image generation. The AR property makes Selftok fundamentally distinct from traditional spatial tokens in the following two key ways: - Selftok offers an elegant and minimalist approach to unify diffusion and AR for vision-language models (VLMs): By representing images with Selftok tokens, we can train a VLM using a purely discrete autoregressive architecture -- like that in LLMs -- without requiring additional modules or training objectives. - We theoretically show that the AR prior satisfies the Bellman equation, whereas the spatial prior does not. Therefore, Selftok supports reinforcement learning (RL) for visual generation with effectiveness comparable to that achieved in LLMs. Besides the AR property, Selftok is also a SoTA tokenizer that achieves a favorable trade-off between high-quality reconstruction and compression rate. We use Selftok to build a pure AR VLM for both visual comprehension and generation tasks. Impressively, without using any text-image training pairs, a simple policy gradient RL working in the visual tokens can significantly boost the visual generation benchmark, surpassing all the existing models by a large margin. Therefore, we believe that Selftok effectively addresses the long-standing challenge that visual tokens cannot support effective RL. When combined with the well-established strengths of RL in LLMs, this brings us one step closer to realizing a truly multimodal LLM. Project Page: https://selftok-team.github.io/report/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Wang, Bohan
Yue, Zhongqi
Zhang, Fengda
Chen, Shuo
Bi, Li'an
Zhang, Junzhe
Song, Xue
Chan, Kennard Yanting
Pan, Jiachun
Wu, Weijia
Zhou, Mingze
Lin, Wang
Pan, Kaihang
Zhang, Saining
Jia, Liyu
Hu, Wentao
Zhao, Wei
Zhang, Hanwang
Computer Vision and Pattern Recognition
We completely discard the conventional spatial prior in image representation and introduce a novel discrete visual tokenizer: Self-consistency Tokenizer (Selftok). At its design core, we compose an autoregressive (AR) prior -- mirroring the causal structure of language -- into visual tokens by using the reverse diffusion process of image generation. The AR property makes Selftok fundamentally distinct from traditional spatial tokens in the following two key ways: - Selftok offers an elegant and minimalist approach to unify diffusion and AR for vision-language models (VLMs): By representing images with Selftok tokens, we can train a VLM using a purely discrete autoregressive architecture -- like that in LLMs -- without requiring additional modules or training objectives. - We theoretically show that the AR prior satisfies the Bellman equation, whereas the spatial prior does not. Therefore, Selftok supports reinforcement learning (RL) for visual generation with effectiveness comparable to that achieved in LLMs. Besides the AR property, Selftok is also a SoTA tokenizer that achieves a favorable trade-off between high-quality reconstruction and compression rate. We use Selftok to build a pure AR VLM for both visual comprehension and generation tasks. Impressively, without using any text-image training pairs, a simple policy gradient RL working in the visual tokens can significantly boost the visual generation benchmark, surpassing all the existing models by a large margin. Therefore, we believe that Selftok effectively addresses the long-standing challenge that visual tokens cannot support effective RL. When combined with the well-established strengths of RL in LLMs, this brings us one step closer to realizing a truly multimodal LLM. Project Page: https://selftok-team.github.io/report/.
title Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.07538