Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Tong, Yujia, Wang, Yuxi, Wan, Yunyang, Zhang, Tian, Dong, Junhao, Yuan, Jingling |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
di: Tong, Yujia, et al.
Pubblicazione: (2026)
di: Tong, Yujia, et al.
Pubblicazione: (2026)
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy
di: Karimi, Hamed, et al.
Pubblicazione: (2026)
di: Karimi, Hamed, et al.
Pubblicazione: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
di: Zhang, Yuxiang, et al.
Pubblicazione: (2025)
di: Zhang, Yuxiang, et al.
Pubblicazione: (2025)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
di: Yi, Ke, et al.
Pubblicazione: (2024)
di: Yi, Ke, et al.
Pubblicazione: (2024)
Conformal Prediction on Quantifying Uncertainty of Dynamic Systems
di: Liang, Aoming, et al.
Pubblicazione: (2024)
di: Liang, Aoming, et al.
Pubblicazione: (2024)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
di: Wu, Fan, et al.
Pubblicazione: (2026)
di: Wu, Fan, et al.
Pubblicazione: (2026)
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
di: Shi, Yang, et al.
Pubblicazione: (2025)
di: Shi, Yang, et al.
Pubblicazione: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
di: Xiong, Boya, et al.
Pubblicazione: (2025)
di: Xiong, Boya, et al.
Pubblicazione: (2025)
English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance
di: Borgersen, Karl Audun, et al.
Pubblicazione: (2025)
di: Borgersen, Karl Audun, et al.
Pubblicazione: (2025)
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
di: Makhov, Denis, et al.
Pubblicazione: (2025)
di: Makhov, Denis, et al.
Pubblicazione: (2025)
Conformal Prediction for Privacy-Preserving Machine Learning
di: Balinsky, Alexander David, et al.
Pubblicazione: (2025)
di: Balinsky, Alexander David, et al.
Pubblicazione: (2025)
Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression
di: Kim, Minjun, et al.
Pubblicazione: (2026)
di: Kim, Minjun, et al.
Pubblicazione: (2026)
Safe Urban Traffic Control via Uncertainty-Aware Conformal Prediction and World-Model Reinforcement Learning
di: Chandra, Joydeep, et al.
Pubblicazione: (2026)
di: Chandra, Joydeep, et al.
Pubblicazione: (2026)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
di: Xia, Junhao, et al.
Pubblicazione: (2025)
di: Xia, Junhao, et al.
Pubblicazione: (2025)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
di: Li, Xuancheng, et al.
Pubblicazione: (2026)
di: Li, Xuancheng, et al.
Pubblicazione: (2026)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
TECP: Token-Entropy Conformal Prediction for LLMs
di: Xu, Beining, et al.
Pubblicazione: (2025)
di: Xu, Beining, et al.
Pubblicazione: (2025)
Robust Machine Unlearning for Quantized Neural Networks via Adaptive Gradient Reweighting with Similar Labels
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
di: Wang, Yanli, et al.
Pubblicazione: (2026)
di: Wang, Yanli, et al.
Pubblicazione: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
Unveil Sources of Uncertainty: Feature Contribution to Conformal Prediction Intervals
di: Idrissi, Marouane Il, et al.
Pubblicazione: (2025)
di: Idrissi, Marouane Il, et al.
Pubblicazione: (2025)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
di: Yang, Xu, et al.
Pubblicazione: (2026)
di: Yang, Xu, et al.
Pubblicazione: (2026)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
di: Lu, Jinming, et al.
Pubblicazione: (2025)
di: Lu, Jinming, et al.
Pubblicazione: (2025)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
di: Gong, Ruihao, et al.
Pubblicazione: (2024)
di: Gong, Ruihao, et al.
Pubblicazione: (2024)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
di: Liu, Lawrence, et al.
Pubblicazione: (2025)
di: Liu, Lawrence, et al.
Pubblicazione: (2025)
RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
di: Zhang, Yuwei, et al.
Pubblicazione: (2024)
di: Zhang, Yuwei, et al.
Pubblicazione: (2024)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
di: Azad, Asif, et al.
Pubblicazione: (2025)
di: Azad, Asif, et al.
Pubblicazione: (2025)
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
di: Jia, Hangyi, et al.
Pubblicazione: (2025)
di: Jia, Hangyi, et al.
Pubblicazione: (2025)
Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
di: Venkatesh, Vinay, et al.
Pubblicazione: (2025)
di: Venkatesh, Vinay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
di: Tong, Yujia, et al.
Pubblicazione: (2025) -
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
di: Zhang, Tian, et al.
Pubblicazione: (2026) -
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
di: Tong, Yujia, et al.
Pubblicazione: (2026) -
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
di: Liu, Xinyu, et al.
Pubblicazione: (2026) -
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)