KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Fei, Liu, Song, Wu, Weiguo, Nie, Shiqiang, Wang, Jinyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Unified Formal Theory on the Logical Limits of Symbol Grounding
von: Liu, Zhangchi
Veröffentlicht: (2025)
von: Liu, Zhangchi
Veröffentlicht: (2025)
Conditional and Modal Reasoning in Large Language Models
von: Holliday, Wesley H., et al.
Veröffentlicht: (2024)
von: Holliday, Wesley H., et al.
Veröffentlicht: (2024)
On the Limits of Learned Importance Scoring for KV Cache Compression
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
von: Abzianidze, Lasha
Veröffentlicht: (2023)
von: Abzianidze, Lasha
Veröffentlicht: (2023)
Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
von: Vossel, Felix, et al.
Veröffentlicht: (2025)
von: Vossel, Felix, et al.
Veröffentlicht: (2025)
TRAWL: Tensor Reduced and Approximated Weights for Large Language Models
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era
von: Brophy, Matthew E.
Veröffentlicht: (2025)
von: Brophy, Matthew E.
Veröffentlicht: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
Robust Hybrid Classical-Quantum Transfer Learning Model for Text Classification Using GPT-Neo 125M with LoRA & SMOTE Enhancement
von: Wishal, Santanam
Veröffentlicht: (2025)
von: Wishal, Santanam
Veröffentlicht: (2025)
FANAL -- Financial Activity News Alerting Language Modeling Framework
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2024)
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
Countermind: A Multi-Layered Security Architecture for Large Language Models
von: Schwarz, Dominik
Veröffentlicht: (2025)
von: Schwarz, Dominik
Veröffentlicht: (2025)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
Encoding Argumentation Frameworks to Propositional Logic Systems
von: Tang, Shuai, et al.
Veröffentlicht: (2025)
von: Tang, Shuai, et al.
Veröffentlicht: (2025)
Are Two Hidden Layers Still Enough for the Physics-Informed Neural Networks?
von: Es'kin, Vasiliy A., et al.
Veröffentlicht: (2024)
von: Es'kin, Vasiliy A., et al.
Veröffentlicht: (2024)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use
von: Santillana, Juan S.
Veröffentlicht: (2026)
von: Santillana, Juan S.
Veröffentlicht: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
Model-Driven Legacy System Modernization at Scale
von: Böhm, Tobias, et al.
Veröffentlicht: (2026)
von: Böhm, Tobias, et al.
Veröffentlicht: (2026)
Neural Tucker Convolutional Network for Water Quality Analysis
von: Si, Hongnan, et al.
Veröffentlicht: (2025)
von: Si, Hongnan, et al.
Veröffentlicht: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
von: Othman, Refat
Veröffentlicht: (2026)
von: Othman, Refat
Veröffentlicht: (2026)
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
von: Xu, Yichun, et al.
Veröffentlicht: (2026)
von: Xu, Yichun, et al.
Veröffentlicht: (2026)
Natural Term Logic
von: Protin, Clarence
Veröffentlicht: (2024)
von: Protin, Clarence
Veröffentlicht: (2024)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
von: Doda, Shravan
Veröffentlicht: (2026)
von: Doda, Shravan
Veröffentlicht: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
Generation of renormalized quadratic coefficient in Landau theory: Implications for specific-heat jump calculations in high-temperature superconductors
von: Claire, Feulefack Ornela, et al.
Veröffentlicht: (2025)
von: Claire, Feulefack Ornela, et al.
Veröffentlicht: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Encoding argumentation frameworks with set attackers to propositional logic systems
von: Tang, Shuai, et al.
Veröffentlicht: (2025)
von: Tang, Shuai, et al.
Veröffentlicht: (2025)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
von: Wang, Haochuan Kevin, et al.
Veröffentlicht: (2026)
von: Wang, Haochuan Kevin, et al.
Veröffentlicht: (2026)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts
von: Elahi, Kazi Toufique, et al.
Veröffentlicht: (2024)
von: Elahi, Kazi Toufique, et al.
Veröffentlicht: (2024)
Correctness is not Faithfulness in RAG Attributions
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
About optimal loss function for training physics-informed neural networks under respecting causality
von: Es'kin, Vasiliy A., et al.
Veröffentlicht: (2023)
von: Es'kin, Vasiliy A., et al.
Veröffentlicht: (2023)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
State-based Modal Logics for Free Choice
von: Aloni, Maria, et al.
Veröffentlicht: (2023)
von: Aloni, Maria, et al.
Veröffentlicht: (2023)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
von: Qi, Jinhu, et al.
Veröffentlicht: (2026)
von: Qi, Jinhu, et al.
Veröffentlicht: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
A Causal Convolutional Low-rank Representation Model for Imputation of Water Quality Data
von: Liao, Xin, et al.
Veröffentlicht: (2025)
von: Liao, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Unified Formal Theory on the Logical Limits of Symbol Grounding
von: Liu, Zhangchi
Veröffentlicht: (2025) -
Conditional and Modal Reasoning in Large Language Models
von: Holliday, Wesley H., et al.
Veröffentlicht: (2024) -
On the Limits of Learned Importance Scoring for KV Cache Compression
von: Steele, Brady
Veröffentlicht: (2026) -
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
von: Abzianidze, Lasha
Veröffentlicht: (2023) -
Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
von: Vossel, Felix, et al.
Veröffentlicht: (2025)