TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuhang, Panda, Priyadarshini |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation
di: Li, Yuhang, et al.
Pubblicazione: (2026)
di: Li, Yuhang, et al.
Pubblicazione: (2026)
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
di: Kim, Jaemin, et al.
Pubblicazione: (2026)
di: Kim, Jaemin, et al.
Pubblicazione: (2026)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
di: Li, Yuhang, et al.
Pubblicazione: (2026)
di: Li, Yuhang, et al.
Pubblicazione: (2026)
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
di: Li, Yuhang, et al.
Pubblicazione: (2025)
di: Li, Yuhang, et al.
Pubblicazione: (2025)
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
di: Yin, Ruokai, et al.
Pubblicazione: (2025)
di: Yin, Ruokai, et al.
Pubblicazione: (2025)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
di: Su, Le, et al.
Pubblicazione: (2026)
di: Su, Le, et al.
Pubblicazione: (2026)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
di: Zhang, Tianao, et al.
Pubblicazione: (2025)
di: Zhang, Tianao, et al.
Pubblicazione: (2025)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
di: Lee, Banseok, et al.
Pubblicazione: (2025)
di: Lee, Banseok, et al.
Pubblicazione: (2025)
Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
di: Zhang, Chenxiang, et al.
Pubblicazione: (2025)
di: Zhang, Chenxiang, et al.
Pubblicazione: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
di: Lee, Deokjae, et al.
Pubblicazione: (2025)
di: Lee, Deokjae, et al.
Pubblicazione: (2025)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
di: Su, Yupeng, et al.
Pubblicazione: (2026)
di: Su, Yupeng, et al.
Pubblicazione: (2026)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
di: Mirzaei, Amir Reza, et al.
Pubblicazione: (2025)
di: Mirzaei, Amir Reza, et al.
Pubblicazione: (2025)
Pushing the Limits of Block Rotations in Post-Training Quantization
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
di: Lee, Donghyun, et al.
Pubblicazione: (2025)
di: Lee, Donghyun, et al.
Pubblicazione: (2025)
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
di: Yan, Xin, et al.
Pubblicazione: (2026)
di: Yan, Xin, et al.
Pubblicazione: (2026)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
di: Li, Ke, et al.
Pubblicazione: (2026)
di: Li, Ke, et al.
Pubblicazione: (2026)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
di: Park, Yeonhong, et al.
Pubblicazione: (2024)
di: Park, Yeonhong, et al.
Pubblicazione: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
GenQ: Quantization in Low Data Regimes with Generative Synthetic Data
di: Li, Yuhang, et al.
Pubblicazione: (2023)
di: Li, Yuhang, et al.
Pubblicazione: (2023)
UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression
di: Zou, Sunan, et al.
Pubblicazione: (2025)
di: Zou, Sunan, et al.
Pubblicazione: (2025)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
di: Zhang, Xi, et al.
Pubblicazione: (2025)
di: Zhang, Xi, et al.
Pubblicazione: (2025)
Why Do Some Inputs Break Low-Bit LLM Quantization?
di: Chang, Ting-Yun, et al.
Pubblicazione: (2025)
di: Chang, Ting-Yun, et al.
Pubblicazione: (2025)
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
di: Lucas, Ryan, et al.
Pubblicazione: (2026)
di: Lucas, Ryan, et al.
Pubblicazione: (2026)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Error Diffusion: Post Training Quantization with Block-Scaled Number Formats for Neural Networks
di: Khodamoradi, Alireza, et al.
Pubblicazione: (2024)
di: Khodamoradi, Alireza, et al.
Pubblicazione: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
ReSpike: Residual Frames-based Hybrid Spiking Neural Networks for Efficient Action Recognition
di: Xiao, Shiting, et al.
Pubblicazione: (2024)
di: Xiao, Shiting, et al.
Pubblicazione: (2024)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
di: Zhang, Jinhao Zhang Yunquan, et al.
Pubblicazione: (2026)
di: Zhang, Jinhao Zhang Yunquan, et al.
Pubblicazione: (2026)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
di: Wang, YiFeng, et al.
Pubblicazione: (2026)
di: Wang, YiFeng, et al.
Pubblicazione: (2026)
Interactions Across Blocks in Post-Training Quantization of Large Language Models
di: Shabanovi, Khasmamad, et al.
Pubblicazione: (2024)
di: Shabanovi, Khasmamad, et al.
Pubblicazione: (2024)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
di: Niu, Muqun, et al.
Pubblicazione: (2024)
di: Niu, Muqun, et al.
Pubblicazione: (2024)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
di: Liao, Baohao, et al.
Pubblicazione: (2024)
di: Liao, Baohao, et al.
Pubblicazione: (2024)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
di: Bang, Wonjun, et al.
Pubblicazione: (2025)
di: Bang, Wonjun, et al.
Pubblicazione: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
di: Xie, Xilong, et al.
Pubblicazione: (2025)
di: Xie, Xilong, et al.
Pubblicazione: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
di: Maskey, Sohir, et al.
Pubblicazione: (2026)
di: Maskey, Sohir, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation
di: Li, Yuhang, et al.
Pubblicazione: (2026) -
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
di: Kim, Jaemin, et al.
Pubblicazione: (2026) -
QuRL: Efficient Reinforcement Learning with Quantized Rollout
di: Li, Yuhang, et al.
Pubblicazione: (2026) -
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
di: Li, Yuhang, et al.
Pubblicazione: (2025) -
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
di: Yin, Ruokai, et al.
Pubblicazione: (2025)