Transformers from Compressed Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alcazar, Juan C. Leon, Soldan, Mattia, Saatialsoruji, Mohammad, Pardo, Alejandro, Itani, Hani, Perez, Juan Camilo, Ghanem, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917048894881792
author Alcazar, Juan C. Leon
Soldan, Mattia
Saatialsoruji, Mohammad
Pardo, Alejandro
Itani, Hani
Perez, Juan Camilo
Ghanem, Bernard
author_facet Alcazar, Juan C. Leon
Soldan, Mattia
Saatialsoruji, Mohammad
Pardo, Alejandro
Itani, Hani
Perez, Juan Camilo
Ghanem, Bernard
contents Compressed file formats are the corner stone of efficient data storage and transmission, yet their potential for representation learning remains largely underexplored. We introduce TEMPEST (TransformErs froM comPressed rEpreSenTations), a method that exploits the inherent byte-stream structure of compressed files to design an effective tokenization and encoding strategy. By leveraging this compact encoding, a standard transformer can directly learn semantic representations from compressed data streams, bypassing the need for raw byte-level processing or full media decoding. Our proposal substantially reduces the number of tokens required for semantic classification, thereby lowering both computational complexity and memory usage. Through extensive experiments across diverse datasets, coding schemes, and modalities, we show that TEMPEST achieves accuracy competitive wit the state-of-the-art while delivering efficiency gains in memory and compute.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformers from Compressed Representations
Alcazar, Juan C. Leon
Soldan, Mattia
Saatialsoruji, Mohammad
Pardo, Alejandro
Itani, Hani
Perez, Juan Camilo
Ghanem, Bernard
Machine Learning
Artificial Intelligence
Compressed file formats are the corner stone of efficient data storage and transmission, yet their potential for representation learning remains largely underexplored. We introduce TEMPEST (TransformErs froM comPressed rEpreSenTations), a method that exploits the inherent byte-stream structure of compressed files to design an effective tokenization and encoding strategy. By leveraging this compact encoding, a standard transformer can directly learn semantic representations from compressed data streams, bypassing the need for raw byte-level processing or full media decoding. Our proposal substantially reduces the number of tokens required for semantic classification, thereby lowering both computational complexity and memory usage. Through extensive experiments across diverse datasets, coding schemes, and modalities, we show that TEMPEST achieves accuracy competitive wit the state-of-the-art while delivering efficiency gains in memory and compute.
title Transformers from Compressed Representations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.23665