Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Badger, Benjamin L., Neligeorge, Matthew |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NanoKnow: How to Know What Your Language Model Knows
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
Language Modeling Is Compression
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
di: Wang, Hanyang, et al.
Pubblicazione: (2026)
di: Wang, Hanyang, et al.
Pubblicazione: (2026)
Memorization-Compression Cycles Improve Generalization
di: Yu, Fangyuan
Pubblicazione: (2025)
di: Yu, Fangyuan
Pubblicazione: (2025)
Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations
di: Halperin, Igor
Pubblicazione: (2025)
di: Halperin, Igor
Pubblicazione: (2025)
Language Model Memory and Memory Models for Language
di: Badger, Benjamin L.
Pubblicazione: (2026)
di: Badger, Benjamin L.
Pubblicazione: (2026)
Learning is Forgetting: LLM Training As Lossy Compression
di: Conklin, Henry C., et al.
Pubblicazione: (2026)
di: Conklin, Henry C., et al.
Pubblicazione: (2026)
Compression Represents Intelligence Linearly
di: Huang, Yuzhen, et al.
Pubblicazione: (2024)
di: Huang, Yuzhen, et al.
Pubblicazione: (2024)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
di: Català, Mar Gonzàlez I, et al.
Pubblicazione: (2026)
di: Català, Mar Gonzàlez I, et al.
Pubblicazione: (2026)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
di: Badger, Benjamin L., et al.
Pubblicazione: (2026)
di: Badger, Benjamin L., et al.
Pubblicazione: (2026)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
di: Zhao, Dongfang
Pubblicazione: (2026)
di: Zhao, Dongfang
Pubblicazione: (2026)
The Information of Large Language Model Geometry
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
A Survey on Large Language Models from Concept to Implementation
di: Wang, Chen, et al.
Pubblicazione: (2024)
di: Wang, Chen, et al.
Pubblicazione: (2024)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
di: Wei, Lai, et al.
Pubblicazione: (2024)
di: Wei, Lai, et al.
Pubblicazione: (2024)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
di: Borzechowski, Florian, et al.
Pubblicazione: (2025)
di: Borzechowski, Florian, et al.
Pubblicazione: (2025)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
di: Jung, Dongwon, et al.
Pubblicazione: (2024)
di: Jung, Dongwon, et al.
Pubblicazione: (2024)
A Training-free Method for LLM Text Attribution
di: Radvand, Tara, et al.
Pubblicazione: (2025)
di: Radvand, Tara, et al.
Pubblicazione: (2025)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
di: Kang, Justin Singh, et al.
Pubblicazione: (2025)
di: Kang, Justin Singh, et al.
Pubblicazione: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
An Information Theoretic Perspective on Agentic System Design
di: He, Shizhe, et al.
Pubblicazione: (2025)
di: He, Shizhe, et al.
Pubblicazione: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
di: Anwar, Usman, et al.
Pubblicazione: (2026)
di: Anwar, Usman, et al.
Pubblicazione: (2026)
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
di: Omidvar, Hamed, et al.
Pubblicazione: (2026)
di: Omidvar, Hamed, et al.
Pubblicazione: (2026)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
di: Tulchinskii, Eduard, et al.
Pubblicazione: (2023)
di: Tulchinskii, Eduard, et al.
Pubblicazione: (2023)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
di: Tsur, Dor, et al.
Pubblicazione: (2025)
di: Tsur, Dor, et al.
Pubblicazione: (2025)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
di: Guo, Yang, et al.
Pubblicazione: (2025)
di: Guo, Yang, et al.
Pubblicazione: (2025)
Entropy-informed Decoding: Adaptive Information-Driven Branching
di: Evans, Benjamin Patrick, et al.
Pubblicazione: (2026)
di: Evans, Benjamin Patrick, et al.
Pubblicazione: (2026)
Agentic Entropy-Balanced Policy Optimization
di: Dong, Guanting, et al.
Pubblicazione: (2025)
di: Dong, Guanting, et al.
Pubblicazione: (2025)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
di: Tohme, Tony, et al.
Pubblicazione: (2023)
di: Tohme, Tony, et al.
Pubblicazione: (2023)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
di: Razzhigaev, Anton, et al.
Pubblicazione: (2023)
di: Razzhigaev, Anton, et al.
Pubblicazione: (2023)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
di: Garg, Sahil, et al.
Pubblicazione: (2024)
di: Garg, Sahil, et al.
Pubblicazione: (2024)
ComMer: a Framework for Compressing and Merging User Data for Personalization
di: Zeldes, Yoel, et al.
Pubblicazione: (2025)
di: Zeldes, Yoel, et al.
Pubblicazione: (2025)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
di: Corallo, Giulio, et al.
Pubblicazione: (2025)
di: Corallo, Giulio, et al.
Pubblicazione: (2025)
Documenti analoghi
-
NanoKnow: How to Know What Your Language Model Knows
di: Gu, Lingwei, et al.
Pubblicazione: (2026) -
Language Modeling Is Compression
di: Delétang, Grégoire, et al.
Pubblicazione: (2023) -
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
di: Wang, Hanyang, et al.
Pubblicazione: (2026) -
Memorization-Compression Cycles Improve Generalization
di: Yu, Fangyuan
Pubblicazione: (2025) -
Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations
di: Halperin, Igor
Pubblicazione: (2025)