Context Channel Capacity: An Information-Theoretic Framework for Understanding Catastrophic Forgetting
Fuente:
arXiv
Saved in:
| Main Author: | Cheng, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Effective Information Theoretic Framework for Channel Pruning
by: Chen, Yihao, et al.
Published: (2024)
by: Chen, Yihao, et al.
Published: (2024)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
An Information-Theoretic Criterion for Efficient Data Synthesis
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
Machine Unlearning via Information Theoretic Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
by: Laakom, Firas, et al.
Published: (2025)
by: Laakom, Firas, et al.
Published: (2025)
Information-Theoretic State Variable Selection for Reinforcement Learning
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
Uncertainty Quantification and Data Efficiency in AI: An Information-Theoretic Perspective
by: Simeone, Osvaldo, et al.
Published: (2025)
by: Simeone, Osvaldo, et al.
Published: (2025)
The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
by: Yang, Jiaming, et al.
Published: (2026)
by: Yang, Jiaming, et al.
Published: (2026)
Broadcast Channel Cooperative Gain: An Operational Interpretation of Partial Information Decomposition
by: Tian, Chao, et al.
Published: (2025)
by: Tian, Chao, et al.
Published: (2025)
Information-Theoretic Policy Pre-Training with Empowerment
by: Schneider, Moritz, et al.
Published: (2025)
by: Schneider, Moritz, et al.
Published: (2025)
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
by: Amaefuna, Theophilus, et al.
Published: (2026)
by: Amaefuna, Theophilus, et al.
Published: (2026)
Probing the Information Theoretical Roots of Spatial Dependence Measures
by: Wang, Zhangyu, et al.
Published: (2024)
by: Wang, Zhangyu, et al.
Published: (2024)
On the Fragility of AI-Based Channel Decoders under Small Channel Perturbations
by: Lei, Haoyu, et al.
Published: (2026)
by: Lei, Haoyu, et al.
Published: (2026)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025)
by: He, Shizhe, et al.
Published: (2025)
Neural Polar Decoders for Deletion Channels
by: Aharoni, Ziv, et al.
Published: (2025)
by: Aharoni, Ziv, et al.
Published: (2025)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
Lost and Found in Translation: Variational Diagnostics for Neural Codebook Channels
by: Hayashi, Yusuke
Published: (2026)
by: Hayashi, Yusuke
Published: (2026)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
by: Omidvar, Hamed, et al.
Published: (2026)
by: Omidvar, Hamed, et al.
Published: (2026)
The Agent Capability Problem: Predicting Solvability Through Information-Theoretic Bounds
by: Lutati, Shahar
Published: (2025)
by: Lutati, Shahar
Published: (2025)
Deep Randomized Distributed Function Computation (DeepRDFC): Neural Distributed Channel Simulation
by: Bergström, Didrik, et al.
Published: (2026)
by: Bergström, Didrik, et al.
Published: (2026)
Continual Learning-Aided Super-Resolution Scheme for Channel Reconstruction and Generalization in OFDM Systems
by: Chen, Jianqiao, et al.
Published: (2025)
by: Chen, Jianqiao, et al.
Published: (2025)
Capacity-Constrained Continual Learning
by: Wen, Zheng, et al.
Published: (2025)
by: Wen, Zheng, et al.
Published: (2025)
Incremental Concept Formation over Visual Images Without Catastrophic Forgetting
by: Barari, Nicki, et al.
Published: (2024)
by: Barari, Nicki, et al.
Published: (2024)
Contextual Control without Memory Growth in a Context-Switching Task
by: Kim, Song-Ju
Published: (2026)
by: Kim, Song-Ju
Published: (2026)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Neural Channel Knowledge Map Assisted Scheduling Optimization of Active IRSs in Multi-User Systems
by: Chen, Xintong, et al.
Published: (2025)
by: Chen, Xintong, et al.
Published: (2025)
Two Birds with One Stone: Multi-Task Semantic Communications Systems over Relay Channel
by: Cao, Yujie, et al.
Published: (2024)
by: Cao, Yujie, et al.
Published: (2024)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
by: Zecchin, Matteo, et al.
Published: (2023)
by: Zecchin, Matteo, et al.
Published: (2023)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Statistical Channel Fingerprint Construction for Massive MIMO: A Unified Tensor Learning Framework
by: Jin, Zhenzhou, et al.
Published: (2026)
by: Jin, Zhenzhou, et al.
Published: (2026)
Catastrophic Forgetting in Kolmogorov-Arnold Networks
by: Rahman, Mohammad Marufur, et al.
Published: (2025)
by: Rahman, Mohammad Marufur, et al.
Published: (2025)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
Information-Theoretic Framework for Understanding Modern Machine-Learning
by: Feder, Meir, et al.
Published: (2025)
by: Feder, Meir, et al.
Published: (2025)
Directed Information $γ$-covering: An Information-Theoretic Framework for Context Engineering
by: Huang, Hai
Published: (2025)
by: Huang, Hai
Published: (2025)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Similar Items
-
An Effective Information Theoretic Framework for Channel Pruning
by: Chen, Yihao, et al.
Published: (2024) -
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026) -
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025) -
An Information-Theoretic Criterion for Efficient Data Synthesis
by: Li, Hanyu, et al.
Published: (2026) -
Machine Unlearning via Information Theoretic Regularization
by: Xu, Shizhou, et al.
Published: (2025)