CODA: Repurposing Continuous VAEs for Discrete Tokenization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Zeyu, Ni, Zanlin, Hua, Yeguo, Deng, Xin, Ma, Xiao, Zhong, Cheng, Huang, Gao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915523347873792
author Liu, Zeyu
Ni, Zanlin
Hua, Yeguo
Deng, Xin
Ma, Xiao
Zhong, Cheng
Huang, Gao
author_facet Liu, Zeyu
Ni, Zanlin
Hua, Yeguo
Deng, Xin
Ma, Xiao
Zhong, Cheng
Huang, Gao
contents Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set of codes. Traditional discrete tokenizers typically learn the two tasks jointly, often leading to unstable training, low codebook utilization, and limited reconstruction quality. In this paper, we introduce \textbf{CODA}(\textbf{CO}ntinuous-to-\textbf{D}iscrete \textbf{A}daptation), a framework that decouples compression and discretization. Instead of training discrete tokenizers from scratch, CODA adapts off-the-shelf continuous VAEs -- already optimized for perceptual compression -- into discrete tokenizers via a carefully designed discretization process. By primarily focusing on discretization, CODA ensures stable and efficient training while retaining the strong visual fidelity of continuous VAEs. Empirically, with $\mathbf{6 \times}$ less training budget than standard VQGAN, our approach achieves a remarkable codebook utilization of 100% and notable reconstruction FID (rFID) of $\mathbf{0.43}$ and $\mathbf{1.34}$ for $8 \times$ and $16 \times$ compression on ImageNet 256$\times$ 256 benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CODA: Repurposing Continuous VAEs for Discrete Tokenization
Liu, Zeyu
Ni, Zanlin
Hua, Yeguo
Deng, Xin
Ma, Xiao
Zhong, Cheng
Huang, Gao
Computer Vision and Pattern Recognition
Artificial Intelligence
Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set of codes. Traditional discrete tokenizers typically learn the two tasks jointly, often leading to unstable training, low codebook utilization, and limited reconstruction quality. In this paper, we introduce \textbf{CODA}(\textbf{CO}ntinuous-to-\textbf{D}iscrete \textbf{A}daptation), a framework that decouples compression and discretization. Instead of training discrete tokenizers from scratch, CODA adapts off-the-shelf continuous VAEs -- already optimized for perceptual compression -- into discrete tokenizers via a carefully designed discretization process. By primarily focusing on discretization, CODA ensures stable and efficient training while retaining the strong visual fidelity of continuous VAEs. Empirically, with $\mathbf{6 \times}$ less training budget than standard VQGAN, our approach achieves a remarkable codebook utilization of 100% and notable reconstruction FID (rFID) of $\mathbf{0.43}$ and $\mathbf{1.34}$ for $8 \times$ and $16 \times$ compression on ImageNet 256$\times$ 256 benchmark.
title CODA: Repurposing Continuous VAEs for Discrete Tokenization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.17760