The Fourth State: Signed-Zero Ternary for Stable LLM Quantization (and More)
Fuente:
arXiv
Salvato in:
| Autore principale: | Uhlmann, Jeffrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unit-Consistent (UC) Adjoint for GSD and Backprop in Deep Learning Applications
di: Uhlmann, Jeffrey
Pubblicazione: (2026)
di: Uhlmann, Jeffrey
Pubblicazione: (2026)
Binary and Ternary Quantization Can Enhance Feature Discrimination
di: Lu, Weizhi, et al.
Pubblicazione: (2025)
di: Lu, Weizhi, et al.
Pubblicazione: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
di: Lipshitz, Baraq, et al.
Pubblicazione: (2025)
di: Lipshitz, Baraq, et al.
Pubblicazione: (2025)
A Geometric Analysis of Sign-Magnitude Asymmetry in a ReLU + RMSNorm Block under Ternary Quantization
di: Dong, Lei
Pubblicazione: (2026)
di: Dong, Lei
Pubblicazione: (2026)
Tequila: Trapping-free Ternary Quantization for Large Language Models
di: Huang, Hong, et al.
Pubblicazione: (2025)
di: Huang, Hong, et al.
Pubblicazione: (2025)
TernaryLLM: Ternarized Large Language Model
di: Chen, Tianqi, et al.
Pubblicazione: (2024)
di: Chen, Tianqi, et al.
Pubblicazione: (2024)
Instance-level Randomization: Toward More Stable LLM Evaluations
di: Li, Yiyang, et al.
Pubblicazione: (2025)
di: Li, Yiyang, et al.
Pubblicazione: (2025)
Quantize What Counts: More for Keys, Less for Values
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
di: Chen, Junyu, et al.
Pubblicazione: (2025)
di: Chen, Junyu, et al.
Pubblicazione: (2025)
ZOQO: Zero-Order Quantized Optimization
di: Bar, Noga, et al.
Pubblicazione: (2025)
di: Bar, Noga, et al.
Pubblicazione: (2025)
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
di: Huang, Hong, et al.
Pubblicazione: (2026)
di: Huang, Hong, et al.
Pubblicazione: (2026)
Towards noise contrastive estimation with soft targets for conditional models
di: Hugger, Johannes, et al.
Pubblicazione: (2024)
di: Hugger, Johannes, et al.
Pubblicazione: (2024)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
di: Zhao, Maosen, et al.
Pubblicazione: (2025)
di: Zhao, Maosen, et al.
Pubblicazione: (2025)
LLaVaOLMoBitnet1B: Ternary LLM goes Multimodal!
di: Sundaram, Jainaveen, et al.
Pubblicazione: (2024)
di: Sundaram, Jainaveen, et al.
Pubblicazione: (2024)
FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
di: Zuo, Fei, et al.
Pubblicazione: (2026)
di: Zuo, Fei, et al.
Pubblicazione: (2026)
Effective Quantization of Muon Optimizer States
di: Gupta, Aman, et al.
Pubblicazione: (2025)
di: Gupta, Aman, et al.
Pubblicazione: (2025)
FPTQuant: Function-Preserving Transforms for LLM Quantization
di: van Breugel, Boris, et al.
Pubblicazione: (2025)
di: van Breugel, Boris, et al.
Pubblicazione: (2025)
KurTail : Kurtosis-based LLM Quantization
di: Akhondzadeh, Mohammad Sadegh, et al.
Pubblicazione: (2025)
di: Akhondzadeh, Mohammad Sadegh, et al.
Pubblicazione: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
di: Ma, Xin, et al.
Pubblicazione: (2026)
di: Ma, Xin, et al.
Pubblicazione: (2026)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
di: Chen, Jiale, et al.
Pubblicazione: (2025)
di: Chen, Jiale, et al.
Pubblicazione: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
di: Panferov, Andrei, et al.
Pubblicazione: (2026)
di: Panferov, Andrei, et al.
Pubblicazione: (2026)
Rethinking Residual Errors in Compensation-based LLM Quantization
di: Li, Shuaiting, et al.
Pubblicazione: (2026)
di: Li, Shuaiting, et al.
Pubblicazione: (2026)
Leech Lattice Vector Quantization for Efficient LLM Compression
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2026)
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2026)
Exploiting LLM Quantization
di: Egashira, Kazuki, et al.
Pubblicazione: (2024)
di: Egashira, Kazuki, et al.
Pubblicazione: (2024)
QERA: an Analytical Framework for Quantization Error Reconstruction
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training
di: Ahn, Myeonghwan, et al.
Pubblicazione: (2025)
di: Ahn, Myeonghwan, et al.
Pubblicazione: (2025)
TTAQ: Towards Stable Post-training Quantization in Continuous Domain Adaptation
di: Xiao, Junrui, et al.
Pubblicazione: (2024)
di: Xiao, Junrui, et al.
Pubblicazione: (2024)
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs
di: Li, Yuhan, et al.
Pubblicazione: (2024)
di: Li, Yuhan, et al.
Pubblicazione: (2024)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
di: van Baalen, Mart, et al.
Pubblicazione: (2024)
di: van Baalen, Mart, et al.
Pubblicazione: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
Mixture of Many Zero-Compute Experts: A High-Rate Quantization Theory Perspective
di: Dar, Yehuda
Pubblicazione: (2025)
di: Dar, Yehuda
Pubblicazione: (2025)
DILEMMA: Joint LLM Quantization and Distributed LLM Inference Over Edge Computing Systems
di: Hosseinzadeh, Minoo, et al.
Pubblicazione: (2025)
di: Hosseinzadeh, Minoo, et al.
Pubblicazione: (2025)
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
di: Ertan, Murat Bilgehan, et al.
Pubblicazione: (2026)
di: Ertan, Murat Bilgehan, et al.
Pubblicazione: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
di: Petrov, Egor, et al.
Pubblicazione: (2025)
di: Petrov, Egor, et al.
Pubblicazione: (2025)
The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
di: Aouad, Ali, et al.
Pubblicazione: (2025)
di: Aouad, Ali, et al.
Pubblicazione: (2025)
TRIX: A More Expressive Model for Zero-shot Domain Transfer in Knowledge Graphs
di: Zhang, Yucheng, et al.
Pubblicazione: (2025)
di: Zhang, Yucheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unit-Consistent (UC) Adjoint for GSD and Backprop in Deep Learning Applications
di: Uhlmann, Jeffrey
Pubblicazione: (2026) -
Binary and Ternary Quantization Can Enhance Feature Discrimination
di: Lu, Weizhi, et al.
Pubblicazione: (2025) -
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
di: Lipshitz, Baraq, et al.
Pubblicazione: (2025) -
A Geometric Analysis of Sign-Magnitude Asymmetry in a ReLU + RMSNorm Block under Ternary Quantization
di: Dong, Lei
Pubblicazione: (2026) -
Tequila: Trapping-free Ternary Quantization for Large Language Models
di: Huang, Hong, et al.
Pubblicazione: (2025)