InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Yang, Wei, Fangyun, He, Tianyu, Zhao, Jinjing, Ni, Zanlin, Liu, Zeyu, Guo, Jiayi, Shi, Lei, Dong, Yue, Chen, Li, Li, Ji, Huang, Gao, Chen, Dong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910219003494400
author Yue, Yang
Wei, Fangyun
He, Tianyu
Zhao, Jinjing
Ni, Zanlin
Liu, Zeyu
Guo, Jiayi
Shi, Lei
Dong, Yue
Chen, Li
Li, Ji
Huang, Gao
Chen, Dong
author_facet Yue, Yang
Wei, Fangyun
He, Tianyu
Zhao, Jinjing
Ni, Zanlin
Liu, Zeyu
Guo, Jiayi
Shi, Lei
Dong, Yue
Chen, Li
Li, Ji
Huang, Gao
Chen, Dong
contents Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content-aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14333
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
Yue, Yang
Wei, Fangyun
He, Tianyu
Zhao, Jinjing
Ni, Zanlin
Liu, Zeyu
Guo, Jiayi
Shi, Lei
Dong, Yue
Chen, Li
Li, Ji
Huang, Gao
Chen, Dong
Computer Vision and Pattern Recognition
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content-aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.
title InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.14333