Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
Fuente:
arXiv
Saved in:
| Main Authors: | Heurtel-Depeiges, David, Ruoss, Anian, Veness, Joel, Genewein, Tim |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025)
by: Genewein, Tim, et al.
Published: (2025)
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024)
by: Grau-Moya, Jordi, et al.
Published: (2024)
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
by: Gilani, Atefeh, et al.
Published: (2026)
by: Gilani, Atefeh, et al.
Published: (2026)
Structure-Aligned Protein Language Model
by: Chen, Can, et al.
Published: (2025)
by: Chen, Can, et al.
Published: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
An Information Criterion for Controlled Disentanglement of Multimodal Data
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
Why is prompting hard? Understanding prompts on binary sequence predictors
by: Wenliang, Li Kevin, et al.
Published: (2025)
by: Wenliang, Li Kevin, et al.
Published: (2025)
Compressing Chemistry Reveals Functional Groups
by: Sharma, Ruben, et al.
Published: (2025)
by: Sharma, Ruben, et al.
Published: (2025)
Distributed and Rate-Adaptive Feature Compression
by: Deshmukh, Aditya, et al.
Published: (2024)
by: Deshmukh, Aditya, et al.
Published: (2024)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
by: Borzechowski, Florian, et al.
Published: (2025)
by: Borzechowski, Florian, et al.
Published: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
by: Narashiman, Swathi Shree, et al.
Published: (2024)
by: Narashiman, Swathi Shree, et al.
Published: (2024)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
by: Ayonrinde, Kola, et al.
Published: (2024)
by: Ayonrinde, Kola, et al.
Published: (2024)
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
by: Niu, Xueyan, et al.
Published: (2026)
by: Niu, Xueyan, et al.
Published: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
by: Kudo, Sota, et al.
Published: (2024)
by: Kudo, Sota, et al.
Published: (2024)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Energy-Efficient Edge Learning via Joint Data Deepening-and-Prefetching
by: Kook, Sujin, et al.
Published: (2024)
by: Kook, Sujin, et al.
Published: (2024)
Informationally Compressive Anonymization: Non-Degrading Sensitive Input Protection for Privacy-Preserving Supervised Machine Learning
by: Samuelson, Jeremy J
Published: (2026)
by: Samuelson, Jeremy J
Published: (2026)
Accelerating Error Correction Code Transformers
by: Levy, Matan, et al.
Published: (2024)
by: Levy, Matan, et al.
Published: (2024)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Pre-training for Recommendation Unlearning
by: Chen, Guoxuan, et al.
Published: (2025)
by: Chen, Guoxuan, et al.
Published: (2025)
Information-Theoretic Policy Pre-Training with Empowerment
by: Schneider, Moritz, et al.
Published: (2025)
by: Schneider, Moritz, et al.
Published: (2025)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
by: Nichani, Arjun, et al.
Published: (2026)
by: Nichani, Arjun, et al.
Published: (2026)
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes
by: Cohen, Shy-el, et al.
Published: (2025)
by: Cohen, Shy-el, et al.
Published: (2025)
Semantic Rate Distortion and Posterior Design: Compute Constraints, Multimodality, and Strategic Inference
by: Akyol, Emrah
Published: (2026)
by: Akyol, Emrah
Published: (2026)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
by: Oomerjee, Adnan, et al.
Published: (2025)
by: Oomerjee, Adnan, et al.
Published: (2025)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
by: Zecchin, Matteo, et al.
Published: (2023)
by: Zecchin, Matteo, et al.
Published: (2023)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
An Information-Theoretic Criterion for Efficient Data Synthesis
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
Partial Information Decomposition for Data Interpretability and Feature Selection
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
Semi-supervised Batch Learning From Logged Data
by: Aminian, Gholamali, et al.
Published: (2022)
by: Aminian, Gholamali, et al.
Published: (2022)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Similar Items
-
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023) -
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024) -
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024) -
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025) -
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024)