Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Zhixuan, Wang, Shaowen, Li, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Simple Model of Inference Scaling Laws
by: Levi, Noam
Published: (2024)
by: Levi, Noam
Published: (2024)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024)
by: Heurtel-Depeiges, David, et al.
Published: (2024)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
by: Khorasani, Arian, et al.
Published: (2026)
by: Khorasani, Arian, et al.
Published: (2026)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Knowledge Graph-Based Explainable and Generalized Zero-Shot Semantic Communications
by: Zhang, Zhaoyu, et al.
Published: (2025)
by: Zhang, Zhaoyu, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Learning During Detection: Continual Learning for Neural OFDM Receivers via DMRS
by: Obeed, Mohanad, et al.
Published: (2026)
by: Obeed, Mohanad, et al.
Published: (2026)
Compressing Chemistry Reveals Functional Groups
by: Sharma, Ruben, et al.
Published: (2025)
by: Sharma, Ruben, et al.
Published: (2025)
Distributed and Rate-Adaptive Feature Compression
by: Deshmukh, Aditya, et al.
Published: (2024)
by: Deshmukh, Aditya, et al.
Published: (2024)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
CSGO: Generalized Optimization for Cold Start in Wireless Collaborative Edge LLM Systems
by: Liu, Xuran, et al.
Published: (2025)
by: Liu, Xuran, et al.
Published: (2025)
Scaling Laws for Online Advertisement Retrieval
by: Wang, Yunli, et al.
Published: (2024)
by: Wang, Yunli, et al.
Published: (2024)
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
by: Gilani, Atefeh, et al.
Published: (2026)
by: Gilani, Atefeh, et al.
Published: (2026)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
by: Borzechowski, Florian, et al.
Published: (2025)
by: Borzechowski, Florian, et al.
Published: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
by: Narashiman, Swathi Shree, et al.
Published: (2024)
by: Narashiman, Swathi Shree, et al.
Published: (2024)
Scaling Laws and Tradeoffs in Recurrent Networks of Expressive Neurons
by: Spieler, Aaron, et al.
Published: (2026)
by: Spieler, Aaron, et al.
Published: (2026)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
by: Ayonrinde, Kola, et al.
Published: (2024)
by: Ayonrinde, Kola, et al.
Published: (2024)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
by: Levy, Jordan, et al.
Published: (2026)
by: Levy, Jordan, et al.
Published: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
by: Kudo, Sota, et al.
Published: (2024)
by: Kudo, Sota, et al.
Published: (2024)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Energy-Efficient Edge Learning via Joint Data Deepening-and-Prefetching
by: Kook, Sujin, et al.
Published: (2024)
by: Kook, Sujin, et al.
Published: (2024)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
by: Badger, Benjamin L., et al.
Published: (2025)
by: Badger, Benjamin L., et al.
Published: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
by: Xu, Zhaoqi, et al.
Published: (2025)
by: Xu, Zhaoqi, et al.
Published: (2025)
Informationally Compressive Anonymization: Non-Degrading Sensitive Input Protection for Privacy-Preserving Supervised Machine Learning
by: Samuelson, Jeremy J
Published: (2026)
by: Samuelson, Jeremy J
Published: (2026)
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
by: Wang, Yikun, et al.
Published: (2026)
by: Wang, Yikun, et al.
Published: (2026)
An Information-Theoretic Criterion for Efficient Data Synthesis
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Context Channel Capacity: An Information-Theoretic Framework for Understanding Catastrophic Forgetting
by: Cheng, Ran
Published: (2026)
by: Cheng, Ran
Published: (2026)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
by: Nichani, Arjun, et al.
Published: (2026)
by: Nichani, Arjun, et al.
Published: (2026)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
An Information Criterion for Controlled Disentanglement of Multimodal Data
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Interpretable Diffusion via Information Decomposition
by: Kong, Xianghao, et al.
Published: (2023)
by: Kong, Xianghao, et al.
Published: (2023)
Neural Channel Knowledge Map Assisted Scheduling Optimization of Active IRSs in Multi-User Systems
by: Chen, Xintong, et al.
Published: (2025)
by: Chen, Xintong, et al.
Published: (2025)
Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
by: Wang, Zhongwei, et al.
Published: (2025)
by: Wang, Zhongwei, et al.
Published: (2025)
Similar Items
-
A Simple Model of Inference Scaling Laws
by: Levi, Noam
Published: (2024) -
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024) -
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026) -
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
by: Wang, Weiqi, et al.
Published: (2026) -
Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
by: Khorasani, Arian, et al.
Published: (2026)