Transformers from Compressed Representations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917048894881792 |
|---|---|
| author | Alcazar, Juan C. Leon Soldan, Mattia Saatialsoruji, Mohammad Pardo, Alejandro Itani, Hani Perez, Juan Camilo Ghanem, Bernard |
| author_facet | Alcazar, Juan C. Leon Soldan, Mattia Saatialsoruji, Mohammad Pardo, Alejandro Itani, Hani Perez, Juan Camilo Ghanem, Bernard |
| contents | Compressed file formats are the corner stone of efficient data storage and transmission, yet their potential for representation learning remains largely underexplored. We introduce TEMPEST (TransformErs froM comPressed rEpreSenTations), a method that exploits the inherent byte-stream structure of compressed files to design an effective tokenization and encoding strategy. By leveraging this compact encoding, a standard transformer can directly learn semantic representations from compressed data streams, bypassing the need for raw byte-level processing or full media decoding. Our proposal substantially reduces the number of tokens required for semantic classification, thereby lowering both computational complexity and memory usage. Through extensive experiments across diverse datasets, coding schemes, and modalities, we show that TEMPEST achieves accuracy competitive wit the state-of-the-art while delivering efficiency gains in memory and compute. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23665 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Transformers from Compressed Representations Alcazar, Juan C. Leon Soldan, Mattia Saatialsoruji, Mohammad Pardo, Alejandro Itani, Hani Perez, Juan Camilo Ghanem, Bernard Machine Learning Artificial Intelligence Compressed file formats are the corner stone of efficient data storage and transmission, yet their potential for representation learning remains largely underexplored. We introduce TEMPEST (TransformErs froM comPressed rEpreSenTations), a method that exploits the inherent byte-stream structure of compressed files to design an effective tokenization and encoding strategy. By leveraging this compact encoding, a standard transformer can directly learn semantic representations from compressed data streams, bypassing the need for raw byte-level processing or full media decoding. Our proposal substantially reduces the number of tokens required for semantic classification, thereby lowering both computational complexity and memory usage. Through extensive experiments across diverse datasets, coding schemes, and modalities, we show that TEMPEST achieves accuracy competitive wit the state-of-the-art while delivering efficiency gains in memory and compute. |
| title | Transformers from Compressed Representations |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2510.23665 |