DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zijian, Huang, Wei, Yu, Yifei, Ren, Tianhe, Wang, Zhongrui, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
RepQuant: Towards Accurate Post-Training Quantization of Large Transformer Models via Scale Reparameterization
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2026)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
von: Xie, Jianhang, et al.
Veröffentlicht: (2025)
von: Xie, Jianhang, et al.
Veröffentlicht: (2025)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
Progressive Binarization with Semi-Structured Pruning for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
Boosting Entropy with Bell Box Quantization
von: Yang, Ningfeng, et al.
Veröffentlicht: (2026)
von: Yang, Ningfeng, et al.
Veröffentlicht: (2026)
Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
von: Tsai, Yu-Lin, et al.
Veröffentlicht: (2023)
von: Tsai, Yu-Lin, et al.
Veröffentlicht: (2023)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
Improving Uncertainty Sampling with Bell Curve Weight Function
von: Chong, Zan-Kai, et al.
Veröffentlicht: (2024)
von: Chong, Zan-Kai, et al.
Veröffentlicht: (2024)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
Unlocking the Pre-Trained Model as a Dual-Alignment Calibrator for Post-Trained LLMs
von: Luo, Beier, et al.
Veröffentlicht: (2026)
von: Luo, Beier, et al.
Veröffentlicht: (2026)
AffineQuant: Affine Transformation Quantization for Large Language Models
von: Ma, Yuexiao, et al.
Veröffentlicht: (2024)
von: Ma, Yuexiao, et al.
Veröffentlicht: (2024)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
von: Chen, Hong, et al.
Veröffentlicht: (2024)
von: Chen, Hong, et al.
Veröffentlicht: (2024)
Breaking Symmetry When Training Transformers
von: Zuo, Chunsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Chunsheng, et al.
Veröffentlicht: (2024)
BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials
von: Xing, Xingrun, et al.
Veröffentlicht: (2023)
von: Xing, Xingrun, et al.
Veröffentlicht: (2023)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
FrameQuant: Flexible Low-Bit Quantization for Transformers
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs
von: Ding, Zhe, et al.
Veröffentlicht: (2026)
von: Ding, Zhe, et al.
Veröffentlicht: (2026)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
Understanding Neural Network Binarization with Forward and Backward Proximal Quantizers
von: Lu, Yiwei, et al.
Veröffentlicht: (2024)
von: Lu, Yiwei, et al.
Veröffentlicht: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
von: Horton, Mark, et al.
Veröffentlicht: (2025)
von: Horton, Mark, et al.
Veröffentlicht: (2025)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
von: Yi, Kai, et al.
Veröffentlicht: (2026)
von: Yi, Kai, et al.
Veröffentlicht: (2026)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026) -
RepQuant: Towards Accurate Post-Training Quantization of Large Transformer Models via Scale Reparameterization
von: Li, Zhikai, et al.
Veröffentlicht: (2024) -
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2026) -
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
von: Xie, Jianhang, et al.
Veröffentlicht: (2025)