Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
Fuente:
arXiv
Saved in:
| Main Authors: | Dickson, Billy, Tiganj, Zoran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
by: Bajaj, Anooshka, et al.
Published: (2026)
by: Bajaj, Anooshka, et al.
Published: (2026)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
by: Kabir, Md Rysul, et al.
Published: (2026)
by: Kabir, Md Rysul, et al.
Published: (2026)
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
by: Bajaj, Anooshka, et al.
Published: (2025)
by: Bajaj, Anooshka, et al.
Published: (2025)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
by: Mistry, Deven Mahesh, et al.
Published: (2025)
by: Mistry, Deven Mahesh, et al.
Published: (2025)
High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
by: Maini, Sahaj Singh, et al.
Published: (2026)
by: Maini, Sahaj Singh, et al.
Published: (2026)
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
by: Huang, Chensen, et al.
Published: (2024)
by: Huang, Chensen, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
by: Kai, Jushi, et al.
Published: (2025)
by: Kai, Jushi, et al.
Published: (2025)
Who Do LLMs Trust? Human Experts Matter More Than Other LLMs
by: Bajaj, Anooshka, et al.
Published: (2026)
by: Bajaj, Anooshka, et al.
Published: (2026)
PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
by: Zhu, Shiyi, et al.
Published: (2023)
by: Zhu, Shiyi, et al.
Published: (2023)
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
by: Jin, Hongye, et al.
Published: (2024)
by: Jin, Hongye, et al.
Published: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
by: Hoang, Duc N. M, et al.
Published: (2023)
by: Hoang, Duc N. M, et al.
Published: (2023)
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation
by: Cao, Bowen, et al.
Published: (2025)
by: Cao, Bowen, et al.
Published: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
GRATH: Gradual Self-Truthifying for Large Language Models
by: Chen, Weixin, et al.
Published: (2024)
by: Chen, Weixin, et al.
Published: (2024)
Long Context Compression with Activation Beacon
by: Zhang, Peitian, et al.
Published: (2024)
by: Zhang, Peitian, et al.
Published: (2024)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
Stacked from One: Multi-Scale Self-Injection for Context Window Extension
by: Han, Wei, et al.
Published: (2026)
by: Han, Wei, et al.
Published: (2026)
Gradually Excavating External Knowledge for Implicit Complex Question Answering
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Enhancing RAG Efficiency with Adaptive Context Compression
by: Guo, Shuyu, et al.
Published: (2025)
by: Guo, Shuyu, et al.
Published: (2025)
Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
by: Wang, Wenshan, et al.
Published: (2024)
by: Wang, Wenshan, et al.
Published: (2024)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024)
by: Dsouza, Amanda, et al.
Published: (2024)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
by: Zhang, Yong, et al.
Published: (2025)
by: Zhang, Yong, et al.
Published: (2025)
Evaluating Zero-Shot Long-Context LLM Compression
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
RL from Teacher-Model Refinement: Gradual Imitation Learning for Machine Translation
by: Lee, Dongyub Jude, et al.
Published: (2025)
by: Lee, Dongyub Jude, et al.
Published: (2025)
Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates
by: Zou, Rui, et al.
Published: (2024)
by: Zou, Rui, et al.
Published: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
by: Zhou, Yiqing, et al.
Published: (2025)
by: Zhou, Yiqing, et al.
Published: (2025)
ACON: Optimizing Context Compression for Long-horizon LLM Agents
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
by: Chari, Vivek, et al.
Published: (2025)
by: Chari, Vivek, et al.
Published: (2025)
AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Transformers Remember First, Forget Last: Dual-Process Interference in LLMs
by: Chattaraj, Sourav, et al.
Published: (2026)
by: Chattaraj, Sourav, et al.
Published: (2026)
YaRN: Efficient Context Window Extension of Large Language Models
by: Peng, Bowen, et al.
Published: (2023)
by: Peng, Bowen, et al.
Published: (2023)
Similar Items
-
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
by: Bajaj, Anooshka, et al.
Published: (2026) -
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
by: Kabir, Md Rysul, et al.
Published: (2026) -
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
by: Bajaj, Anooshka, et al.
Published: (2025) -
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
by: Mistry, Deven Mahesh, et al.
Published: (2025) -
High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
by: Maini, Sahaj Singh, et al.
Published: (2026)