STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Yuhang, Yang, Wenzheng, Chen, Yujie, Jin, Xiangqi, Zhang, Yaojie, Huang, Siteng, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Kwai-STaR: Transform LLMs into State-Transition Reasoners
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
Adaptive KV-Cache Compression without Manually Setting Budget
von: Tang, Chenxia, et al.
Veröffentlicht: (2025)
von: Tang, Chenxia, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
STaR-SQL: Self-Taught Reasoner for Text-to-SQL
von: He, Mingqian, et al.
Veröffentlicht: (2025)
von: He, Mingqian, et al.
Veröffentlicht: (2025)
HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
von: Guo, Shaoxiong, et al.
Veröffentlicht: (2025)
von: Guo, Shaoxiong, et al.
Veröffentlicht: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Lean-STaR: Learning to Interleave Thinking and Proving
von: Lin, Haohan, et al.
Veröffentlicht: (2024)
von: Lin, Haohan, et al.
Veröffentlicht: (2024)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
von: Andukuri, Chinmaya, et al.
Veröffentlicht: (2024)
von: Andukuri, Chinmaya, et al.
Veröffentlicht: (2024)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency Constraints
von: Yang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Yang, Xiaohang, et al.
Veröffentlicht: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
von: Rehg, Isaac
Veröffentlicht: (2024)
von: Rehg, Isaac
Veröffentlicht: (2024)
Graph-Guided Adaptive Channel Elimination for KV Cache Compression
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
von: Li, Xuelin, et al.
Veröffentlicht: (2025) -
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025) -
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025) -
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025) -
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)