JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914186245701632 |
|---|---|
| author | Ioannides, Georgios Constantinou, Christos Chadha, Aman Elkins, Aaron Pang, Linsey Shwartz-Ziv, Ravid LeCun, Yann |
| author_facet | Ioannides, Georgios Constantinou, Christos Chadha, Aman Elkins, Aaron Pang, Linsey Shwartz-Ziv, Ravid LeCun, Yann |
| contents | We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_07168 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention Ioannides, Georgios Constantinou, Christos Chadha, Aman Elkins, Aaron Pang, Linsey Shwartz-Ziv, Ravid LeCun, Yann Sound Artificial Intelligence Machine Learning Audio and Speech Processing We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs. |
| title | JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention |
| topic | Sound Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.07168 |