Breaking the KV Cache Bottleneck: Fan Duality Model Achieves O(1) Decode Memory with Superior Associative Recall
Fuente:
arXiv
Salvato in:
| Autore principale: | Fan, Yasong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators
di: Sridhar, Anupama, et al.
Pubblicazione: (2026)
di: Sridhar, Anupama, et al.
Pubblicazione: (2026)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
di: Tan, Yifan, et al.
Pubblicazione: (2024)
di: Tan, Yifan, et al.
Pubblicazione: (2024)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
di: Liu, Sihao, et al.
Pubblicazione: (2026)
di: Liu, Sihao, et al.
Pubblicazione: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
di: Li, Kunjun, et al.
Pubblicazione: (2025)
di: Li, Kunjun, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
di: Roy, Sourjya, et al.
Pubblicazione: (2025)
di: Roy, Sourjya, et al.
Pubblicazione: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
di: Liu, Andy Zeyi, et al.
Pubblicazione: (2026)
di: Liu, Andy Zeyi, et al.
Pubblicazione: (2026)
From TLinFormer to TConstFormer: The Leap to Constant-Time Transformer Attention: Achieving O(1) Computation and O(1) KV Cache during Autoregressive Inference
di: Tang, Zhongpan
Pubblicazione: (2025)
di: Tang, Zhongpan
Pubblicazione: (2025)
Training Transformers for KV Cache Compressibility
di: Gelberg, Yoav, et al.
Pubblicazione: (2026)
di: Gelberg, Yoav, et al.
Pubblicazione: (2026)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
di: Vejendla, Harshil
Pubblicazione: (2025)
di: Vejendla, Harshil
Pubblicazione: (2025)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
di: Gokhale, Sai, et al.
Pubblicazione: (2025)
di: Gokhale, Sai, et al.
Pubblicazione: (2025)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
di: Mao, Yuzhen, et al.
Pubblicazione: (2026)
di: Mao, Yuzhen, et al.
Pubblicazione: (2026)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
The Pitfalls of KV Cache Compression
di: Chen, Alex, et al.
Pubblicazione: (2025)
di: Chen, Alex, et al.
Pubblicazione: (2025)
Compute Or Load KV Cache? Why Not Both?
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
Breaking the Attention Bottleneck
di: Hilsenbek, Kalle
Pubblicazione: (2024)
di: Hilsenbek, Kalle
Pubblicazione: (2024)
Understanding Factual Recall in Transformers via Associative Memories
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
di: Yang, Bin, et al.
Pubblicazione: (2025)
di: Yang, Bin, et al.
Pubblicazione: (2025)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
di: Dong, Zican, et al.
Pubblicazione: (2026)
di: Dong, Zican, et al.
Pubblicazione: (2026)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
Breaking Symmetry Bottlenecks in GNN Readouts
di: Talhi, Mouad, et al.
Pubblicazione: (2026)
di: Talhi, Mouad, et al.
Pubblicazione: (2026)
FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
di: Takbir, Nazmul, et al.
Pubblicazione: (2025)
di: Takbir, Nazmul, et al.
Pubblicazione: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
di: Hendria, Willy Fitra
Pubblicazione: (2026)
di: Hendria, Willy Fitra
Pubblicazione: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
di: Liu, Akide, et al.
Pubblicazione: (2024)
di: Liu, Akide, et al.
Pubblicazione: (2024)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators
di: Sridhar, Anupama, et al.
Pubblicazione: (2026) -
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
di: Tomar, Aditya, et al.
Pubblicazione: (2025) -
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024) -
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025) -
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
di: Tan, Yifan, et al.
Pubblicazione: (2024)