Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Yujia, Wang, Yuxi, Wan, Yunyang, Zhang, Tian, Dong, Junhao, Yuan, Jingling |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
by: Tong, Yujia, et al.
Published: (2026)
by: Tong, Yujia, et al.
Published: (2026)
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy
by: Karimi, Hamed, et al.
Published: (2026)
by: Karimi, Hamed, et al.
Published: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
by: Cho, Yoonjun, et al.
Published: (2026)
by: Cho, Yoonjun, et al.
Published: (2026)
Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
by: Yi, Ke, et al.
Published: (2024)
by: Yi, Ke, et al.
Published: (2024)
Conformal Prediction on Quantifying Uncertainty of Dynamic Systems
by: Liang, Aoming, et al.
Published: (2024)
by: Liang, Aoming, et al.
Published: (2024)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
by: Xiong, Boya, et al.
Published: (2025)
by: Xiong, Boya, et al.
Published: (2025)
English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance
by: Borgersen, Karl Audun, et al.
Published: (2025)
by: Borgersen, Karl Audun, et al.
Published: (2025)
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
by: Makhov, Denis, et al.
Published: (2025)
by: Makhov, Denis, et al.
Published: (2025)
Conformal Prediction for Privacy-Preserving Machine Learning
by: Balinsky, Alexander David, et al.
Published: (2025)
by: Balinsky, Alexander David, et al.
Published: (2025)
Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression
by: Kim, Minjun, et al.
Published: (2026)
by: Kim, Minjun, et al.
Published: (2026)
Safe Urban Traffic Control via Uncertainty-Aware Conformal Prediction and World-Model Reinforcement Learning
by: Chandra, Joydeep, et al.
Published: (2026)
by: Chandra, Joydeep, et al.
Published: (2026)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
by: Liu, Yijun, et al.
Published: (2024)
by: Liu, Yijun, et al.
Published: (2024)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
by: Xia, Junhao, et al.
Published: (2025)
by: Xia, Junhao, et al.
Published: (2025)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
TECP: Token-Entropy Conformal Prediction for LLMs
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
Robust Machine Unlearning for Quantized Neural Networks via Adaptive Gradient Reweighting with Similar Labels
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
by: Wang, Yanli, et al.
Published: (2026)
by: Wang, Yanli, et al.
Published: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Unveil Sources of Uncertainty: Feature Contribution to Conformal Prediction Intervals
by: Idrissi, Marouane Il, et al.
Published: (2025)
by: Idrissi, Marouane Il, et al.
Published: (2025)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
by: Gong, Ruihao, et al.
Published: (2024)
by: Gong, Ruihao, et al.
Published: (2024)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
by: Liu, Lawrence, et al.
Published: (2025)
by: Liu, Lawrence, et al.
Published: (2025)
RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
by: Zhang, Yuwei, et al.
Published: (2024)
by: Zhang, Yuwei, et al.
Published: (2024)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
by: Azad, Asif, et al.
Published: (2025)
by: Azad, Asif, et al.
Published: (2025)
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
by: Jia, Hangyi, et al.
Published: (2025)
by: Jia, Hangyi, et al.
Published: (2025)
Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
by: Venkatesh, Vinay, et al.
Published: (2025)
by: Venkatesh, Vinay, et al.
Published: (2025)
Similar Items
-
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
by: Tong, Yujia, et al.
Published: (2025) -
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
by: Zhang, Tian, et al.
Published: (2026) -
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
by: Tong, Yujia, et al.
Published: (2026) -
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
by: Liu, Xinyu, et al.
Published: (2026) -
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
by: Chiang, Hung-Yueh, et al.
Published: (2025)