Quantize What Counts: More for Keys, Less for Values
Fuente:
arXiv
Salvato in:
| Autori principali: | Hariri, Mohsen, Luo, Alan, Chen, Weicong, Zhong, Shaochen, Zhang, Tianyi, Wang, Qifan, Hu, Xia, Han, Xiaotian, Chaudhary, Vipin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
di: Yang, Wang, et al.
Pubblicazione: (2025)
di: Yang, Wang, et al.
Pubblicazione: (2025)
$K^4$: Online Log Anomaly Detection Via Unsupervised Typicality Learning
di: Chen, Weicong, et al.
Pubblicazione: (2025)
di: Chen, Weicong, et al.
Pubblicazione: (2025)
Scorio.jl: A Julia package for ranking stochastic responses
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
Ranking Reasoning LLMs under Test-Time Scaling
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
di: Wang, Shouren, et al.
Pubblicazione: (2025)
di: Wang, Shouren, et al.
Pubblicazione: (2025)
SELF: Self-Extend the Context Length With Logistic Growth Function
di: Dang, Phat Thanh, et al.
Pubblicazione: (2025)
di: Dang, Phat Thanh, et al.
Pubblicazione: (2025)
Reliability-Gated Source Anchoring for Continual Test-Time Adaptation
di: Singh, Vikash, et al.
Pubblicazione: (2026)
di: Singh, Vikash, et al.
Pubblicazione: (2026)
Medical Image Spatial Grounding with Semantic Sampling
di: Yu, Andrew Seohwan, et al.
Pubblicazione: (2026)
di: Yu, Andrew Seohwan, et al.
Pubblicazione: (2026)
Thinking Preference Optimization
di: Yang, Wang, et al.
Pubblicazione: (2025)
di: Yang, Wang, et al.
Pubblicazione: (2025)
Robust Ultra-High-Dimensional Variable Selection With Correlated Structure Using Group Testing
di: Guo, Wanru, et al.
Pubblicazione: (2026)
di: Guo, Wanru, et al.
Pubblicazione: (2026)
CausalGuard: Conformal Inference under Graph Uncertainty
di: Singh, Vikash, et al.
Pubblicazione: (2026)
di: Singh, Vikash, et al.
Pubblicazione: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
di: Liu, Zirui, et al.
Pubblicazione: (2024)
di: Liu, Zirui, et al.
Pubblicazione: (2024)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
di: Yang, Wang, et al.
Pubblicazione: (2025)
di: Yang, Wang, et al.
Pubblicazione: (2025)
Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
di: Liu, Zirui, et al.
Pubblicazione: (2023)
di: Liu, Zirui, et al.
Pubblicazione: (2023)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
di: Guo, Yiju, et al.
Pubblicazione: (2026)
di: Guo, Yiju, et al.
Pubblicazione: (2026)
AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
di: Luo, Feng, et al.
Pubblicazione: (2025)
di: Luo, Feng, et al.
Pubblicazione: (2025)
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
di: Han, Xiaotian, et al.
Pubblicazione: (2023)
di: Han, Xiaotian, et al.
Pubblicazione: (2023)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
Trust The Typical
di: Ganguly, Debargha, et al.
Pubblicazione: (2026)
di: Ganguly, Debargha, et al.
Pubblicazione: (2026)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
di: Wang, Guanchu, et al.
Pubblicazione: (2024)
di: Wang, Guanchu, et al.
Pubblicazione: (2024)
Novel adaptation of video segmentation to 3D MRI: efficient zero-shot knee segmentation with SAM2
di: Yu, Andrew Seohwan, et al.
Pubblicazione: (2024)
di: Yu, Andrew Seohwan, et al.
Pubblicazione: (2024)
When Less is More: The LLM Scaling Paradox in Context Compression
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
di: Zhang, Michael S., et al.
Pubblicazione: (2025)
di: Zhang, Michael S., et al.
Pubblicazione: (2025)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
di: Yang, Wang, et al.
Pubblicazione: (2025)
di: Yang, Wang, et al.
Pubblicazione: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
di: Schulte, David, et al.
Pubblicazione: (2024)
di: Schulte, David, et al.
Pubblicazione: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
di: Deng, Xun, et al.
Pubblicazione: (2025)
di: Deng, Xun, et al.
Pubblicazione: (2025)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
di: Yang, Wang, et al.
Pubblicazione: (2026)
di: Yang, Wang, et al.
Pubblicazione: (2026)
Forget Less, Generalize More: Unifying Temporal and Structural Adaptation for Dynamic Graphs
di: Chang, Qian, et al.
Pubblicazione: (2026)
di: Chang, Qian, et al.
Pubblicazione: (2026)
Less Random, More Private: What is the Optimal Subsampling Scheme for DP-SGD?
di: Dong, Andy, et al.
Pubblicazione: (2026)
di: Dong, Andy, et al.
Pubblicazione: (2026)
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
di: Hu, Chengming, et al.
Pubblicazione: (2023)
di: Hu, Chengming, et al.
Pubblicazione: (2023)
Visual Concept Networks: A Graph-Based Approach to Detecting Anomalous Data in Deep Neural Networks
di: Ganguly, Debargha, et al.
Pubblicazione: (2024)
di: Ganguly, Debargha, et al.
Pubblicazione: (2024)
When Less is More: On the Value of "Co-training" for Semi-Supervised Software Defect Predictors
di: Majumder, Suvodeep, et al.
Pubblicazione: (2022)
di: Majumder, Suvodeep, et al.
Pubblicazione: (2022)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
di: Shen, Yuhao, et al.
Pubblicazione: (2026)
di: Shen, Yuhao, et al.
Pubblicazione: (2026)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
di: Duanmu, Haojie, et al.
Pubblicazione: (2024)
di: Duanmu, Haojie, et al.
Pubblicazione: (2024)
Less is More: Pseudo-Label Filtering for Continual Test-Time Adaptation
di: Tan, Jiayao, et al.
Pubblicazione: (2024)
di: Tan, Jiayao, et al.
Pubblicazione: (2024)
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
di: Yang, Wang, et al.
Pubblicazione: (2026)
di: Yang, Wang, et al.
Pubblicazione: (2026)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
di: Liu, Hongyi, et al.
Pubblicazione: (2024)
di: Liu, Hongyi, et al.
Pubblicazione: (2024)
Transformer Multivariate Forecasting: Less is More?
di: Xu, Jingjing, et al.
Pubblicazione: (2023)
di: Xu, Jingjing, et al.
Pubblicazione: (2023)
Documenti analoghi
-
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025) -
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
di: Yang, Wang, et al.
Pubblicazione: (2025) -
$K^4$: Online Log Anomaly Detection Via Unsupervised Typicality Learning
di: Chen, Weicong, et al.
Pubblicazione: (2025) -
Scorio.jl: A Julia package for ranking stochastic responses
di: Hariri, Mohsen, et al.
Pubblicazione: (2026) -
Ranking Reasoning LLMs under Test-Time Scaling
di: Hariri, Mohsen, et al.
Pubblicazione: (2026)