Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiong, Boya, Wang, Shuo, Ge, Weifeng, Chen, Guanhua, Chen, Yun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
di: Li, Yixia, et al.
Pubblicazione: (2024)
di: Li, Yixia, et al.
Pubblicazione: (2024)
ERC-SVD: Error-Controlled SVD for Large Language Model Compression
di: Bai, Haolei, et al.
Pubblicazione: (2025)
di: Bai, Haolei, et al.
Pubblicazione: (2025)
On the Compressibility of Quantized Large Language Models
di: Mao, Yu, et al.
Pubblicazione: (2024)
di: Mao, Yu, et al.
Pubblicazione: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
di: Ball, Thomas, et al.
Pubblicazione: (2024)
di: Ball, Thomas, et al.
Pubblicazione: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
In-context Autoencoder for Context Compression in a Large Language Model
di: Ge, Tao, et al.
Pubblicazione: (2023)
di: Ge, Tao, et al.
Pubblicazione: (2023)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
di: Zhao, Taibiao, et al.
Pubblicazione: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
di: Salfati, Samuel
Pubblicazione: (2026)
di: Salfati, Samuel
Pubblicazione: (2026)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
di: Yu, Erxin, et al.
Pubblicazione: (2025)
di: Yu, Erxin, et al.
Pubblicazione: (2025)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
di: Zou, Minghui, et al.
Pubblicazione: (2024)
di: Zou, Minghui, et al.
Pubblicazione: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
di: Zhang, Hailin, et al.
Pubblicazione: (2024)
di: Zhang, Hailin, et al.
Pubblicazione: (2024)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
di: Liu, Chi, et al.
Pubblicazione: (2026)
di: Liu, Chi, et al.
Pubblicazione: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
di: Yang, Dongquan, et al.
Pubblicazione: (2025)
di: Yang, Dongquan, et al.
Pubblicazione: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
di: Li, He, et al.
Pubblicazione: (2024)
di: Li, He, et al.
Pubblicazione: (2024)
Retrieval Enhanced Feedback via In-context Neural Error-book
di: Hyun, Jongyeop, et al.
Pubblicazione: (2025)
di: Hyun, Jongyeop, et al.
Pubblicazione: (2025)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
di: Li, Ziniu, et al.
Pubblicazione: (2025)
di: Li, Ziniu, et al.
Pubblicazione: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
di: Wang, Boshi, et al.
Pubblicazione: (2024)
di: Wang, Boshi, et al.
Pubblicazione: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
di: Gong, Ruihao, et al.
Pubblicazione: (2024)
di: Gong, Ruihao, et al.
Pubblicazione: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
di: Bozoukov, Matthew, et al.
Pubblicazione: (2025)
di: Bozoukov, Matthew, et al.
Pubblicazione: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
di: Yu, Wenhao, et al.
Pubblicazione: (2025)
di: Yu, Wenhao, et al.
Pubblicazione: (2025)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
di: Paglieri, Davide, et al.
Pubblicazione: (2024)
di: Paglieri, Davide, et al.
Pubblicazione: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
Adapting LLMs for Efficient Context Processing through Soft Prompt Compression
di: Wang, Cangqing, et al.
Pubblicazione: (2024)
di: Wang, Cangqing, et al.
Pubblicazione: (2024)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
di: Chen, Ruishuo, et al.
Pubblicazione: (2026)
di: Chen, Ruishuo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
di: Li, Yixia, et al.
Pubblicazione: (2024) -
ERC-SVD: Error-Controlled SVD for Large Language Model Compression
di: Bai, Haolei, et al.
Pubblicazione: (2025) -
On the Compressibility of Quantized Large Language Models
di: Mao, Yu, et al.
Pubblicazione: (2024) -
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025) -
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025)