Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Yicheng, Zhong, Zhizhou, Zhang, Jun, Yang, Qin, Jin, XiTai, Qin, Ying, Luo, Wenhan, Mao, Shuiyang, Liu, Wei, Li, Huan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
von: Rehg, Isaac
Veröffentlicht: (2024)
von: Rehg, Isaac
Veröffentlicht: (2024)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling
von: Cederlund, Jonathan, et al.
Veröffentlicht: (2026)
von: Cederlund, Jonathan, et al.
Veröffentlicht: (2026)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026) -
HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
von: Zeng, Bowen, et al.
Veröffentlicht: (2026) -
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025) -
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
von: Luo, Jiayi, et al.
Veröffentlicht: (2026) -
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)