Memorization-Compression Cycles Improve Generalization
Fuente:
arXiv
Saved in:
| Main Author: | Yu, Fangyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
by: Badger, Benjamin L., et al.
Published: (2025)
by: Badger, Benjamin L., et al.
Published: (2025)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
by: Anwar, Usman, et al.
Published: (2026)
by: Anwar, Usman, et al.
Published: (2026)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
by: Lee, Jin Hwa, et al.
Published: (2024)
by: Lee, Jin Hwa, et al.
Published: (2024)
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
by: Jung, Dongwon, et al.
Published: (2024)
by: Jung, Dongwon, et al.
Published: (2024)
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025)
by: Radvand, Tara, et al.
Published: (2025)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025)
by: Wieser, Frederico, et al.
Published: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025)
by: He, Shizhe, et al.
Published: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
Optimal Quantization for Matrix Multiplication
by: Ordentlich, Or, et al.
Published: (2024)
by: Ordentlich, Or, et al.
Published: (2024)
A Survey on Large Language Models from Concept to Implementation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
by: Omidvar, Hamed, et al.
Published: (2026)
by: Omidvar, Hamed, et al.
Published: (2026)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
by: Guo, Yang, et al.
Published: (2025)
by: Guo, Yang, et al.
Published: (2025)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026)
by: Garg, Nikhil, et al.
Published: (2026)
OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions
by: Huang, Yu-Shin, et al.
Published: (2024)
by: Huang, Yu-Shin, et al.
Published: (2024)
Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations
by: Halperin, Igor
Published: (2025)
by: Halperin, Igor
Published: (2025)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
by: Garg, Sahil, et al.
Published: (2024)
by: Garg, Sahil, et al.
Published: (2024)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
by: Zhao, Dongfang
Published: (2026)
by: Zhao, Dongfang
Published: (2026)
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
ComMer: a Framework for Compressing and Merging User Data for Personalization
by: Zeldes, Yoel, et al.
Published: (2025)
by: Zeldes, Yoel, et al.
Published: (2025)
Memorization in In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2024)
by: Golchin, Shahriar, et al.
Published: (2024)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
by: Sriram, Aniruddh, et al.
Published: (2024)
by: Sriram, Aniruddh, et al.
Published: (2024)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025)
by: Corallo, Giulio, et al.
Published: (2025)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
by: Lim, Woosang, et al.
Published: (2025)
by: Lim, Woosang, et al.
Published: (2025)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
by: Zhang, Xuechen, et al.
Published: (2025)
by: Zhang, Xuechen, et al.
Published: (2025)
SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors
by: Trukhina, Natalia, et al.
Published: (2026)
by: Trukhina, Natalia, et al.
Published: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
by: Barbulescu, George-Octavian, et al.
Published: (2024)
by: Barbulescu, George-Octavian, et al.
Published: (2024)
Mitigating Memorization In Language Models
by: Sakarvadia, Mansi, et al.
Published: (2024)
by: Sakarvadia, Mansi, et al.
Published: (2024)
Similar Items
-
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
by: Badger, Benjamin L., et al.
Published: (2025) -
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023) -
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024) -
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026) -
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
by: Anwar, Usman, et al.
Published: (2026)