GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xuan, Zhang, Zhongyu, Huang, Yuge, Mi, Yuxi, Mu, Guodong, Ding, Shouhong, Wang, Jun, Guo, Rizen, Zhou, Shuigeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908704181321728
author Zhao, Xuan
Zhang, Zhongyu
Huang, Yuge
Mi, Yuxi
Mu, Guodong
Ding, Shouhong
Wang, Jun
Guo, Rizen
Zhou, Shuigeng
author_facet Zhao, Xuan
Zhang, Zhongyu
Huang, Yuge
Mi, Yuxi
Mu, Guodong
Ding, Shouhong
Wang, Jun
Guo, Rizen
Zhou, Shuigeng
contents Existing state-of-the-art image tokenization methods leverage diverse semantic features from pre-trained vision models for additional supervision, to expand the distribution of latent representations and thereby improve the quality of image reconstruction and generation. These methods employ a locally supervised approach for semantic supervision, which limits the uniformity of semantic distribution. However, VA-VAE proves that a more uniform feature distribution yields better generation performance. In this work, we introduce a Global Perspective Tokenizer (GloTok), which utilizes global relational information to model a more uniform semantic distribution of tokenized features. Specifically, a codebook-wise histogram relation learning method is proposed to transfer the semantics, which are modeled by pre-trained models on the entire dataset, to the semantic codebook. Then, we design a residual learning module that recovers the fine-grained details to minimize the reconstruction error caused by quantization. Through the above design, GloTok delivers more uniformly distributed semantic latent representations, which facilitates the training of autoregressive (AR) models for generating high-quality images without requiring direct access to pre-trained models during the training process. Experiments on the standard ImageNet-1k benchmark clearly show that our proposed method achieves state-of-the-art reconstruction performance and generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
Zhao, Xuan
Zhang, Zhongyu
Huang, Yuge
Mi, Yuxi
Mu, Guodong
Ding, Shouhong
Wang, Jun
Guo, Rizen
Zhou, Shuigeng
Computer Vision and Pattern Recognition
Existing state-of-the-art image tokenization methods leverage diverse semantic features from pre-trained vision models for additional supervision, to expand the distribution of latent representations and thereby improve the quality of image reconstruction and generation. These methods employ a locally supervised approach for semantic supervision, which limits the uniformity of semantic distribution. However, VA-VAE proves that a more uniform feature distribution yields better generation performance. In this work, we introduce a Global Perspective Tokenizer (GloTok), which utilizes global relational information to model a more uniform semantic distribution of tokenized features. Specifically, a codebook-wise histogram relation learning method is proposed to transfer the semantics, which are modeled by pre-trained models on the entire dataset, to the semantic codebook. Then, we design a residual learning module that recovers the fine-grained details to minimize the reconstruction error caused by quantization. Through the above design, GloTok delivers more uniformly distributed semantic latent representations, which facilitates the training of autoregressive (AR) models for generating high-quality images without requiring direct access to pre-trained models during the training process. Experiments on the standard ImageNet-1k benchmark clearly show that our proposed method achieves state-of-the-art reconstruction performance and generation quality.
title GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.14184