Discrete JEPA: Learning Discrete Token Representations without Reconstruction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Baek, Junyeob, Lee, Hosung, Hoang, Christopher, Ren, Mengye, Ahn, Sungjin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911018489217024
author Baek, Junyeob
Lee, Hosung
Hoang, Christopher
Ren, Mengye
Ahn, Sungjin
author_facet Baek, Junyeob
Lee, Hosung
Hoang, Christopher
Ren, Mengye
Ahn, Sungjin
contents The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant limitations in tasks requiring symbolic abstraction and logical reasoning capabilities essential for systematic inference. To address this challenge, we propose Discrete-JEPA, extending the latent predictive coding framework with semantic tokenization and novel complementary objectives to create robust tokenization for symbolic reasoning tasks. Discrete-JEPA dramatically outperforms baselines on visual symbolic prediction tasks, while striking visual evidence reveals the spontaneous emergence of deliberate systematic patterns within the learned semantic token space. Though an initial model, our approach promises a significant impact for advancing Symbolic world modeling and planning capabilities in artificial intelligence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discrete JEPA: Learning Discrete Token Representations without Reconstruction
Baek, Junyeob
Lee, Hosung
Hoang, Christopher
Ren, Mengye
Ahn, Sungjin
Computer Vision and Pattern Recognition
The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant limitations in tasks requiring symbolic abstraction and logical reasoning capabilities essential for systematic inference. To address this challenge, we propose Discrete-JEPA, extending the latent predictive coding framework with semantic tokenization and novel complementary objectives to create robust tokenization for symbolic reasoning tasks. Discrete-JEPA dramatically outperforms baselines on visual symbolic prediction tasks, while striking visual evidence reveals the spontaneous emergence of deliberate systematic patterns within the learned semantic token space. Though an initial model, our approach promises a significant impact for advancing Symbolic world modeling and planning capabilities in artificial intelligence systems.
title Discrete JEPA: Learning Discrete Token Representations without Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.14373