Gespeichert in:
| Hauptverfasser: | Mao, Weian, Lin, Xi, Huang, Wei, Xie, Yuxin, Fu, Tianfu, Zhuang, Bohan, Han, Song, Chen, Yukang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.04921 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)
von: Su, Yi, et al.
Veröffentlicht: (2026)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
von: Li, Xiaolong, et al.
Veröffentlicht: (2025)
von: Li, Xiaolong, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
von: Rehg, Isaac
Veröffentlicht: (2024)
von: Rehg, Isaac
Veröffentlicht: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
von: Zhuang, Shuhan, et al.
Veröffentlicht: (2025)
von: Zhuang, Shuhan, et al.
Veröffentlicht: (2025)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
LongVLM: Efficient Long Video Understanding via Large Language Models
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
WeGeFT: Weight-Generative Fine-Tuning for Multi-Faceted Efficient Adaptation of Large Models
von: Savadikar, Chinmay, et al.
Veröffentlicht: (2023)
von: Savadikar, Chinmay, et al.
Veröffentlicht: (2023)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
von: Huang, Kai, et al.
Veröffentlicht: (2025)
von: Huang, Kai, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
von: Li, Guihong, et al.
Veröffentlicht: (2025)
von: Li, Guihong, et al.
Veröffentlicht: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
von: Cao, Yukang, et al.
Veröffentlicht: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
von: Song, Jiwon, et al.
Veröffentlicht: (2025)
von: Song, Jiwon, et al.
Veröffentlicht: (2025)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
von: Chen, Yukang, et al.
Veröffentlicht: (2026) -
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025) -
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026) -
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024) -
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)