Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Oomerjee, Adnan, Fountas, Zafeirios, Bou-Ammar, Haitham, Wang, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
by: Fountas, Zafeirios, et al.
Published: (2026)
by: Fountas, Zafeirios, et al.
Published: (2026)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025)
by: Wieser, Frederico, et al.
Published: (2025)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
by: Benfeghoul, Martin, et al.
Published: (2025)
by: Benfeghoul, Martin, et al.
Published: (2025)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
by: Hazard, Hugo, et al.
Published: (2025)
by: Hazard, Hugo, et al.
Published: (2025)
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024)
by: Fountas, Zafeirios, et al.
Published: (2024)
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
by: Yang, Jiaming, et al.
Published: (2026)
by: Yang, Jiaming, et al.
Published: (2026)
Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs
by: Ma, Julian, et al.
Published: (2025)
by: Ma, Julian, et al.
Published: (2025)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
The Alignment Bottleneck
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
The Causal Information Bottleneck and Optimal Causal Variable Abstractions
by: Simoes, Francisco N. F. Q., et al.
Published: (2024)
by: Simoes, Francisco N. F. Q., et al.
Published: (2024)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025)
by: Corallo, Giulio, et al.
Published: (2025)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
by: Kudo, Sota, et al.
Published: (2024)
by: Kudo, Sota, et al.
Published: (2024)
Latent Representation and Simulation of Markov Processes via Time-Lagged Information Bottleneck
by: Federici, Marco, et al.
Published: (2023)
by: Federici, Marco, et al.
Published: (2023)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
by: Benfeghoul, Martin, et al.
Published: (2024)
by: Benfeghoul, Martin, et al.
Published: (2024)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
by: Chen, Kaiwen, et al.
Published: (2025)
by: Chen, Kaiwen, et al.
Published: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
by: Nguyen, Tu, et al.
Published: (2026)
by: Nguyen, Tu, et al.
Published: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
by: Yang, Bin, et al.
Published: (2025)
by: Yang, Bin, et al.
Published: (2025)
Accelerating Error Correction Code Transformers
by: Levy, Matan, et al.
Published: (2024)
by: Levy, Matan, et al.
Published: (2024)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes
by: Cohen, Shy-el, et al.
Published: (2025)
by: Cohen, Shy-el, et al.
Published: (2025)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
by: Gao, Yifei, et al.
Published: (2026)
by: Gao, Yifei, et al.
Published: (2026)
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
by: Zecchin, Matteo, et al.
Published: (2023)
by: Zecchin, Matteo, et al.
Published: (2023)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024)
by: Heurtel-Depeiges, David, et al.
Published: (2024)
A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
by: Urbina-Rodriguez, Pedro, et al.
Published: (2026)
by: Urbina-Rodriguez, Pedro, et al.
Published: (2026)
Similar Items
-
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
by: Fountas, Zafeirios, et al.
Published: (2026) -
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025) -
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
by: Benfeghoul, Martin, et al.
Published: (2025) -
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
by: Hazard, Hugo, et al.
Published: (2025) -
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024)