SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Huanxuan, Xu, Yixing, He, Shizhu, Li, Guanchen, Yin, Xuanwu, Li, Dong, Barsoum, Emad, Zhao, Jun, Liu, Kang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
by: Li, Guanchen, et al.
Published: (2025)
by: Li, Guanchen, et al.
Published: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026)
by: Li, Zekai, et al.
Published: (2026)
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
by: Joo, Donghyeon, et al.
Published: (2025)
by: Joo, Donghyeon, et al.
Published: (2025)
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
by: Li, Jinze, et al.
Published: (2026)
by: Li, Jinze, et al.
Published: (2026)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge Enhancement
by: Tan, Yuqiao, et al.
Published: (2025)
by: Tan, Yuqiao, et al.
Published: (2025)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
by: Joshi, Vinay, et al.
Published: (2025)
by: Joshi, Vinay, et al.
Published: (2025)
Theory-optimal Quantization Based on Flatness
by: Huang, Xiusheng, et al.
Published: (2026)
by: Huang, Xiusheng, et al.
Published: (2026)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
by: Li, Guanchen, et al.
Published: (2024)
by: Li, Guanchen, et al.
Published: (2024)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks
by: Liao, Huanxuan, et al.
Published: (2024)
by: Liao, Huanxuan, et al.
Published: (2024)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
$\textit{SKIntern}$: Internalizing Symbolic Knowledge for Distilling Better CoT Capabilities into Small Language Models
by: Liao, Huanxuan, et al.
Published: (2024)
by: Liao, Huanxuan, et al.
Published: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
by: Zhang, Te, et al.
Published: (2025)
by: Zhang, Te, et al.
Published: (2025)
Awakening Augmented Generation: Learning to Awaken Internal Knowledge of Large Language Models for Question Answering
by: Liao, Huanxuan, et al.
Published: (2024)
by: Liao, Huanxuan, et al.
Published: (2024)
ThinK: Thinner Key Cache by Query-Driven Pruning
by: Xu, Yuhui, et al.
Published: (2024)
by: Xu, Yuhui, et al.
Published: (2024)
Towards Threshold-Free KV Cache Pruning
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
by: Tu, Dezhan, et al.
Published: (2024)
by: Tu, Dezhan, et al.
Published: (2024)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
by: Zhu, Haowei, et al.
Published: (2026)
by: Zhu, Haowei, et al.
Published: (2026)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026)
by: Wu, Guanlong, et al.
Published: (2026)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
by: Ke, Wenjin, et al.
Published: (2025)
by: Ke, Wenjin, et al.
Published: (2025)
From Instance Training to Instruction Learning: Task Adapters Generation from Instructions
by: Liao, Huanxuan, et al.
Published: (2024)
by: Liao, Huanxuan, et al.
Published: (2024)
Shuttle Between the Instructions and the Parameters of Large Language Models
by: Sun, Wangtao, et al.
Published: (2025)
by: Sun, Wangtao, et al.
Published: (2025)
Query2Triple: Unified Query Encoding for Answering Diverse Complex Queries over Knowledge Graphs
by: Xu, Yao, et al.
Published: (2023)
by: Xu, Yao, et al.
Published: (2023)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
by: Zhang, Yike, et al.
Published: (2025)
by: Zhang, Yike, et al.
Published: (2025)
Analytical Queries for Unstructured Data
by: Kang, Daniel
Published: (2025)
by: Kang, Daniel
Published: (2025)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
by: Jegou, Simon, et al.
Published: (2026)
by: Jegou, Simon, et al.
Published: (2026)
CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
by: Chen, Chuangtao, et al.
Published: (2026)
by: Chen, Chuangtao, et al.
Published: (2026)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
Similar Items
-
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
by: Li, Guanchen, et al.
Published: (2025) -
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026) -
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
by: Joo, Donghyeon, et al.
Published: (2025) -
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025) -
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)