Adaptive Length Image Tokenization via Recurrent Allocation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Duggal, Shivam, Isola, Phillip, Torralba, Antonio, Freeman, William T.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929575225720832
author Duggal, Shivam
Isola, Phillip
Torralba, Antonio
Freeman, William T.
author_facet Duggal, Shivam
Isola, Phillip
Torralba, Antonio
Freeman, William T.
contents Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intelligence - and even large language models - which allocate varying representational capacities based on entropy, context and familiarity. Inspired by this, we propose an approach to learn variable-length token representations for 2D images. Our encoder-decoder architecture recursively processes 2D image tokens, distilling them into 1D latent tokens over multiple iterations of recurrent rollouts. Each iteration refines the 2D tokens, updates the existing 1D latent tokens, and adaptively increases representational capacity by adding new tokens. This enables compression of images into a variable number of tokens, ranging from 32 to 256. We validate our tokenizer using reconstruction loss and FID metrics, demonstrating that token count aligns with image entropy, familiarity and downstream task requirements. Recurrent token processing with increasing representational capacity in each iteration shows signs of token specialization, revealing potential for object / part discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02393
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Length Image Tokenization via Recurrent Allocation
Duggal, Shivam
Isola, Phillip
Torralba, Antonio
Freeman, William T.
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intelligence - and even large language models - which allocate varying representational capacities based on entropy, context and familiarity. Inspired by this, we propose an approach to learn variable-length token representations for 2D images. Our encoder-decoder architecture recursively processes 2D image tokens, distilling them into 1D latent tokens over multiple iterations of recurrent rollouts. Each iteration refines the 2D tokens, updates the existing 1D latent tokens, and adaptively increases representational capacity by adding new tokens. This enables compression of images into a variable number of tokens, ranging from 32 to 256. We validate our tokenizer using reconstruction loss and FID metrics, demonstrating that token count aligns with image entropy, familiarity and downstream task requirements. Recurrent token processing with increasing representational capacity in each iteration shows signs of token specialization, revealing potential for object / part discovery.
title Adaptive Length Image Tokenization via Recurrent Allocation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2411.02393