Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zihan, Tang, Cheng, Gong, Lei, Li, Cheng, Wang, Chao, wang, teng, Lou, Wenqi, Zhou, Xuehai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
por: Ramachandran, Akshat, et al.
Publicado: (2025)
por: Ramachandran, Akshat, et al.
Publicado: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
por: Chen, Chuangtao, et al.
Publicado: (2026)
por: Chen, Chuangtao, et al.
Publicado: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025)
por: Hu, Jie, et al.
Publicado: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
por: Jiang, Zhonghua, et al.
Publicado: (2025)
por: Jiang, Zhonghua, et al.
Publicado: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
Reducing End-to-End Latency of Cause-Effect Chains with Shared Cache Analysis
por: Zhu, Yixuan, et al.
Publicado: (2026)
por: Zhu, Yixuan, et al.
Publicado: (2026)
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
por: Tan, Yifan, et al.
Publicado: (2024)
por: Tan, Yifan, et al.
Publicado: (2024)
PiKV: KV Cache Management System for Mixture of Experts
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
por: Zhang, Rongzhi, et al.
Publicado: (2024)
por: Zhang, Rongzhi, et al.
Publicado: (2024)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
por: Liao, Bingli, et al.
Publicado: (2024)
por: Liao, Bingli, et al.
Publicado: (2024)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
por: Wu, Wenbo, et al.
Publicado: (2025)
por: Wu, Wenbo, et al.
Publicado: (2025)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
por: Qi, Yanlin, et al.
Publicado: (2026)
por: Qi, Yanlin, et al.
Publicado: (2026)
QAQ: Quality Adaptive Quantization for LLM KV Cache
por: Dong, Shichen, et al.
Publicado: (2024)
por: Dong, Shichen, et al.
Publicado: (2024)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026)
por: Chen, Jian, et al.
Publicado: (2026)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
por: Wang, Yuxuan, et al.
Publicado: (2026)
por: Wang, Yuxuan, et al.
Publicado: (2026)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
por: Liu, Hongyao, et al.
Publicado: (2026)
por: Liu, Hongyao, et al.
Publicado: (2026)
CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
por: Wang, Rui, et al.
Publicado: (2025)
por: Wang, Rui, et al.
Publicado: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
por: Cai, Zefan, et al.
Publicado: (2025)
por: Cai, Zefan, et al.
Publicado: (2025)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
por: Mao, Yuzhen, et al.
Publicado: (2026)
por: Mao, Yuzhen, et al.
Publicado: (2026)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
por: Tang, Zihan, et al.
Publicado: (2026)
por: Tang, Zihan, et al.
Publicado: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
por: Zou, Jing, et al.
Publicado: (2026)
por: Zou, Jing, et al.
Publicado: (2026)
Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
por: Liu, Xiaoran, et al.
Publicado: (2025)
por: Liu, Xiaoran, et al.
Publicado: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
por: Hao, Jitai, et al.
Publicado: (2026)
por: Hao, Jitai, et al.
Publicado: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
por: Xiong, Yi, et al.
Publicado: (2024)
por: Xiong, Yi, et al.
Publicado: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
por: Xu, Boxun, et al.
Publicado: (2025)
por: Xu, Boxun, et al.
Publicado: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
por: Zeng, Wenxuan, et al.
Publicado: (2025)
por: Zeng, Wenxuan, et al.
Publicado: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
por: Lei, Jianlong, et al.
Publicado: (2026)
por: Lei, Jianlong, et al.
Publicado: (2026)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
por: Zhou, Bowen, et al.
Publicado: (2026)
por: Zhou, Bowen, et al.
Publicado: (2026)
KV Cache Steering for Controlling Frozen LLMs
por: Belitsky, Max, et al.
Publicado: (2025)
por: Belitsky, Max, et al.
Publicado: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
por: Sun, Qiheng, et al.
Publicado: (2025)
por: Sun, Qiheng, et al.
Publicado: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
por: Liu, Minghui, et al.
Publicado: (2025)
por: Liu, Minghui, et al.
Publicado: (2025)
Ejemplares similares
-
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
por: Ramachandran, Akshat, et al.
Publicado: (2025) -
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
por: Chen, Chuangtao, et al.
Publicado: (2026) -
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025) -
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)