Discrete JEPA: Learning Discrete Token Representations without Reconstruction
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911018489217024 |
|---|---|
| author | Baek, Junyeob Lee, Hosung Hoang, Christopher Ren, Mengye Ahn, Sungjin |
| author_facet | Baek, Junyeob Lee, Hosung Hoang, Christopher Ren, Mengye Ahn, Sungjin |
| contents | The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant limitations in tasks requiring symbolic abstraction and logical reasoning capabilities essential for systematic inference. To address this challenge, we propose Discrete-JEPA, extending the latent predictive coding framework with semantic tokenization and novel complementary objectives to create robust tokenization for symbolic reasoning tasks. Discrete-JEPA dramatically outperforms baselines on visual symbolic prediction tasks, while striking visual evidence reveals the spontaneous emergence of deliberate systematic patterns within the learned semantic token space. Though an initial model, our approach promises a significant impact for advancing Symbolic world modeling and planning capabilities in artificial intelligence systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_14373 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Discrete JEPA: Learning Discrete Token Representations without Reconstruction Baek, Junyeob Lee, Hosung Hoang, Christopher Ren, Mengye Ahn, Sungjin Computer Vision and Pattern Recognition The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant limitations in tasks requiring symbolic abstraction and logical reasoning capabilities essential for systematic inference. To address this challenge, we propose Discrete-JEPA, extending the latent predictive coding framework with semantic tokenization and novel complementary objectives to create robust tokenization for symbolic reasoning tasks. Discrete-JEPA dramatically outperforms baselines on visual symbolic prediction tasks, while striking visual evidence reveals the spontaneous emergence of deliberate systematic patterns within the learned semantic token space. Though an initial model, our approach promises a significant impact for advancing Symbolic world modeling and planning capabilities in artificial intelligence systems. |
| title | Discrete JEPA: Learning Discrete Token Representations without Reconstruction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.14373 |