SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jialong, Wang, Zhenglin, Zhang, Linhai, Lai, Yilong, He, Yulan, Zhou, Deyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
by: Lai, Yilong, et al.
Published: (2025)
by: Lai, Yilong, et al.
Published: (2025)
PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
by: Zhang, Linhai, et al.
Published: (2025)
by: Zhang, Linhai, et al.
Published: (2025)
SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
by: Wang, Zhenglin, et al.
Published: (2024)
by: Wang, Zhenglin, et al.
Published: (2024)
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024)
by: Zhang, Congzhi, et al.
Published: (2024)
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation
by: Zhang, Linhai, et al.
Published: (2025)
by: Zhang, Linhai, et al.
Published: (2025)
DINER: Debiasing Aspect-based Sentiment Analysis with Multi-variable Causal Inference
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models
by: Zhang, Linhai, et al.
Published: (2024)
by: Zhang, Linhai, et al.
Published: (2024)
Large, Small or Both: A Novel Data Augmentation Framework Based on Language Models for Debiasing Opinion Summarization
by: Zhang, Yanyue, et al.
Published: (2024)
by: Zhang, Yanyue, et al.
Published: (2024)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
RGAR: Recurrence Generation-augmented Retrieval for Factual-aware Medical Question Answering
by: Liang, Sichu, et al.
Published: (2025)
by: Liang, Sichu, et al.
Published: (2025)
AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment
by: Lai, Yilong, et al.
Published: (2024)
by: Lai, Yilong, et al.
Published: (2024)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
by: Wang, Zhenglin, et al.
Published: (2025)
by: Wang, Zhenglin, et al.
Published: (2025)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
by: Peng, Junjie, et al.
Published: (2026)
by: Peng, Junjie, et al.
Published: (2026)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
by: Chen, Hong, et al.
Published: (2026)
by: Chen, Hong, et al.
Published: (2026)
Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024)
by: Zhang, Congzhi, et al.
Published: (2024)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
by: He, Ziwei, et al.
Published: (2025)
by: He, Ziwei, et al.
Published: (2025)
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Depictor : Topic‐Guided Opinion Summarization for Product Reviews With Dual‐Perspective Topic Modeling
by: Yanyue Zhang, et al.
Published: (2026)
by: Yanyue Zhang, et al.
Published: (2026)
Rehearse With User: Personalized Opinion Summarization via Role-Playing based on Large Language Models
by: Zhang, Yanyue, et al.
Published: (2025)
by: Zhang, Yanyue, et al.
Published: (2025)
SCOPE: A Generative Approach for LLM Prompt Compression
by: Zhang, Tinghui, et al.
Published: (2025)
by: Zhang, Tinghui, et al.
Published: (2025)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
by: Javidnia, Neusha, et al.
Published: (2025)
by: Javidnia, Neusha, et al.
Published: (2025)
CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
by: Lai, Yilong, et al.
Published: (2025)
by: Lai, Yilong, et al.
Published: (2025)
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
by: Liang, Sichu, et al.
Published: (2026)
by: Liang, Sichu, et al.
Published: (2026)
Fine-grainedly Synthesize Streaming Data Based On Large Language Models With Graph Structure Understanding For Data Sparsity
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty
by: Zhong, Meizhi, et al.
Published: (2024)
by: Zhong, Meizhi, et al.
Published: (2024)
MA$^{2}$P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion
by: Zhang, Dingyi, et al.
Published: (2026)
by: Zhang, Dingyi, et al.
Published: (2026)
Persuasion Should be Double-Blind: A Multi-Domain Dialogue Dataset With Faithfulness Based on Causal Theory of Mind
by: Zhang, Dingyi, et al.
Published: (2025)
by: Zhang, Dingyi, et al.
Published: (2025)
SynGraph: A Dynamic Graph-LLM Synthesis Framework for Sparse Streaming User Sentiment Modeling
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
by: Zhao, Shangziqi, et al.
Published: (2025)
by: Zhao, Shangziqi, et al.
Published: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
Learning to Evict from Key-Value Cache
by: Moschella, Luca, et al.
Published: (2026)
by: Moschella, Luca, et al.
Published: (2026)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
by: Zhou, Xiabin, et al.
Published: (2024)
by: Zhou, Xiabin, et al.
Published: (2024)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
by: Yuan, Jian, et al.
Published: (2025)
by: Yuan, Jian, et al.
Published: (2025)
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
by: Zhang, Hongxuan, et al.
Published: (2024)
by: Zhang, Hongxuan, et al.
Published: (2024)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
by: Shi, Zhiyuan, et al.
Published: (2026)
by: Shi, Zhiyuan, et al.
Published: (2026)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
by: Cui, Wanyun, et al.
Published: (2025)
by: Cui, Wanyun, et al.
Published: (2025)
Similar Items
-
AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
by: Lai, Yilong, et al.
Published: (2025) -
PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
by: Zhang, Linhai, et al.
Published: (2025) -
SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
by: Wang, Zhenglin, et al.
Published: (2024) -
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024) -
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation
by: Zhang, Linhai, et al.
Published: (2025)