Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boroujeni, Sayed Pedram Haeri, Mehrabi, Niloufar, Woods, Patrick, Hillesheim, Gabriel, Razi, Abolfazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025)
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025)
From Talking Words to Sharing Thoughts: Scalable Multi-LLM Aggregation via Structured Message Passing
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2026)
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2026)
Adaptive Data Transport Mechanism for UAV Surveillance Missions in Lossy Environments
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2024)
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2024)
Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025)
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025)
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication
von: Sarlak, Ahmad, et al.
Veröffentlicht: (2024)
von: Sarlak, Ahmad, et al.
Veröffentlicht: (2024)
Meta-Adaptive Beam Search Planning for Transformer-Based Reinforcement Learning Control of UAVs with Overhead Manipulators under Flight Disturbances
von: Alzorgan, Hazim, et al.
Veröffentlicht: (2026)
von: Alzorgan, Hazim, et al.
Veröffentlicht: (2026)
FLAME Diffuser: Wildfire Image Synthesis using Mask Guided Diffusion
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Integrating Random Regret Minimization-Based Discrete Choice Models with Mixed Integer Linear Programming for Revenue Optimization
von: Talebi, Amirreza, et al.
Veröffentlicht: (2024)
von: Talebi, Amirreza, et al.
Veröffentlicht: (2024)
Opinion Dynamics in Social Multiplex Networks with Mono and Bi-directional Interactions in the Presence of Leaders
von: Talebi, Amirreza, et al.
Veröffentlicht: (2024)
von: Talebi, Amirreza, et al.
Veröffentlicht: (2024)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
Enhancing Graph Neural Networks in Large-scale Traffic Incident Analysis with Concurrency Hypothesis
von: Chen, Xiwen, et al.
Veröffentlicht: (2024)
von: Chen, Xiwen, et al.
Veröffentlicht: (2024)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents
von: Vahidi-Moghaddam, Amin, et al.
Veröffentlicht: (2025)
von: Vahidi-Moghaddam, Amin, et al.
Veröffentlicht: (2025)
Driving Towards Inclusion: A Systematic Review of AI-powered Accessibility Enhancements for People with Disability in Autonomous Vehicles
von: Bastola, Ashish, et al.
Veröffentlicht: (2024)
von: Bastola, Ashish, et al.
Veröffentlicht: (2024)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
von: Xu, Boxun, et al.
Veröffentlicht: (2025)
Make Your LVLM KV Cache More Lightweight
von: Chen, Xihao, et al.
Veröffentlicht: (2026)
von: Chen, Xihao, et al.
Veröffentlicht: (2026)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
Comparative Analysis of Patch Attack on VLM-Based Autonomous Driving Architectures
von: Fernandez, David, et al.
Veröffentlicht: (2026)
von: Fernandez, David, et al.
Veröffentlicht: (2026)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
Imaging Signal Recovery Using Neural Network Priors Under Uncertain Forward Model Parameters
von: Chen, Xiwen, et al.
Veröffentlicht: (2024)
von: Chen, Xiwen, et al.
Veröffentlicht: (2024)
AtomDiffuser: Time-Aware Degradation Modeling for Drift and Beam Damage in STEM Imaging
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)
von: Tao, Wei, et al.
Veröffentlicht: (2026)
Don't let the information slip away
von: Li, Taozhe, et al.
Veröffentlicht: (2026)
von: Li, Taozhe, et al.
Veröffentlicht: (2026)
Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
Vision Transformers Don't Need Trained Registers
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
Don't Pause! Every prediction matters in a streaming video
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2026)
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2026)
Don't Look into the Dark: Latent Codes for Pluralistic Image Inpainting
von: Chen, Haiwei, et al.
Veröffentlicht: (2024)
von: Chen, Haiwei, et al.
Veröffentlicht: (2024)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Don't Fear Peculiar Activation Functions: EUAF and Beyond
von: Wang, Qianchao, et al.
Veröffentlicht: (2024)
von: Wang, Qianchao, et al.
Veröffentlicht: (2024)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025) -
From Talking Words to Sharing Thoughts: Scalable Multi-LLM Aggregation via Structured Message Passing
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2026) -
Adaptive Data Transport Mechanism for UAV Surveillance Missions in Lossy Environments
von: Mehrabi, Niloufar, et al.
Veröffentlicht: (2024) -
Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection
von: Boroujeni, Sayed Pedram Haeri, et al.
Veröffentlicht: (2025) -
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication
von: Sarlak, Ahmad, et al.
Veröffentlicht: (2024)