Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Ziran, Lv, Youru, Lin, Mingbao, Guo, Hang, Zhang, Zeren, Zou, Danping, Lin, Weiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling
von: Cederlund, Jonathan, et al.
Veröffentlicht: (2026)
von: Cederlund, Jonathan, et al.
Veröffentlicht: (2026)
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
von: Guo, Hang, et al.
Veröffentlicht: (2025)
von: Guo, Hang, et al.
Veröffentlicht: (2025)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
von: Huang, Kai, et al.
Veröffentlicht: (2025)
von: Huang, Kai, et al.
Veröffentlicht: (2025)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Efficient Autoregressive Video Diffusion with Dummy Head
von: Guo, Hang, et al.
Veröffentlicht: (2026)
von: Guo, Hang, et al.
Veröffentlicht: (2026)
Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
von: Xie, Rui, et al.
Veröffentlicht: (2024)
von: Xie, Rui, et al.
Veröffentlicht: (2024)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)
von: He, Zhihao, et al.
Veröffentlicht: (2026)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
Parallelized Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
von: Hu, Lianyu, et al.
Veröffentlicht: (2025)
von: Hu, Lianyu, et al.
Veröffentlicht: (2025)
Visual Implicit Autoregressive Modeling
von: Jiang, Pengfei, et al.
Veröffentlicht: (2026)
von: Jiang, Pengfei, et al.
Veröffentlicht: (2026)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
von: Zuo, Dexin, et al.
Veröffentlicht: (2025)
von: Zuo, Dexin, et al.
Veröffentlicht: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Efficient Diffusion Transformers
von: Chen, Guantao, et al.
Veröffentlicht: (2026)
von: Chen, Guantao, et al.
Veröffentlicht: (2026)
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
von: Qin, Ziran, et al.
Veröffentlicht: (2025) -
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026) -
HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling
von: Cederlund, Jonathan, et al.
Veröffentlicht: (2026) -
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025) -
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
von: Xu, Boxun, et al.
Veröffentlicht: (2025)