LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Gordon, Ofir, Dikstein, Lior, Netzer, Arnon, Achituve, Idan, Habi, Hai Victor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Data Generation for Hardware-Friendly Post-Training Quantization
di: Dikstein, Lior, et al.
Pubblicazione: (2024)
di: Dikstein, Lior, et al.
Pubblicazione: (2024)
EPTQ: Enhanced Post-Training Quantization via Hessian-guided Network-wise Optimization
di: Gordon, Ofir, et al.
Pubblicazione: (2023)
di: Gordon, Ofir, et al.
Pubblicazione: (2023)
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
di: Gordon, Ofir, et al.
Pubblicazione: (2025)
di: Gordon, Ofir, et al.
Pubblicazione: (2025)
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
di: Achituve, Idan, et al.
Pubblicazione: (2025)
di: Achituve, Idan, et al.
Pubblicazione: (2025)
Efficient Image Restoration via Latent Consistency Flow Matching
di: Cohen, Elad, et al.
Pubblicazione: (2025)
di: Cohen, Elad, et al.
Pubblicazione: (2025)
Bayesian Uncertainty for Gradient Aggregation in Multi-Task Learning
di: Achituve, Idan, et al.
Pubblicazione: (2024)
di: Achituve, Idan, et al.
Pubblicazione: (2024)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
De-Confusing Pseudo-Labels in Source-Free Domain Adaptation
di: Diamant, Idit, et al.
Pubblicazione: (2024)
di: Diamant, Idit, et al.
Pubblicazione: (2024)
Refusal in LLMs is an Affine Function
di: Marshall, Thomas, et al.
Pubblicazione: (2024)
di: Marshall, Thomas, et al.
Pubblicazione: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
di: Li, Zekai, et al.
Pubblicazione: (2026)
di: Li, Zekai, et al.
Pubblicazione: (2026)
Diverse Sampling in Diffusion Models with Marginal Preserving Particle Guidance
di: Vinograd, Gal, et al.
Pubblicazione: (2026)
di: Vinograd, Gal, et al.
Pubblicazione: (2026)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
di: Aksenov, Yaroslav, et al.
Pubblicazione: (2024)
di: Aksenov, Yaroslav, et al.
Pubblicazione: (2024)
Learning When to Quit in Sales Conversations
di: Manzoor, Emaad, et al.
Pubblicazione: (2025)
di: Manzoor, Emaad, et al.
Pubblicazione: (2025)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
di: Arbel, Iftach, et al.
Pubblicazione: (2024)
di: Arbel, Iftach, et al.
Pubblicazione: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
di: Lee, Changhun, et al.
Pubblicazione: (2024)
di: Lee, Changhun, et al.
Pubblicazione: (2024)
How Does Quantization Affect Multilingual LLMs?
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
di: Fasoli, Andrea, et al.
Pubblicazione: (2026)
di: Fasoli, Andrea, et al.
Pubblicazione: (2026)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
di: Choi, Euntae, et al.
Pubblicazione: (2025)
di: Choi, Euntae, et al.
Pubblicazione: (2025)
LESA: Learnable LLM Layer Scaling-Up
di: Yang, Yifei, et al.
Pubblicazione: (2025)
di: Yang, Yifei, et al.
Pubblicazione: (2025)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
di: Mazza, Arnon, et al.
Pubblicazione: (2026)
di: Mazza, Arnon, et al.
Pubblicazione: (2026)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
di: Kiruluta, Andrew
Pubblicazione: (2025)
di: Kiruluta, Andrew
Pubblicazione: (2025)
LQER: Low-Rank Quantization Error Reconstruction for LLMs
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
di: Li, Zhikai, et al.
Pubblicazione: (2023)
di: Li, Zhikai, et al.
Pubblicazione: (2023)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
di: Shoham, Ofir Ben, et al.
Pubblicazione: (2024)
di: Shoham, Ofir Ben, et al.
Pubblicazione: (2024)
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers
di: Neshaei, Seyed Parsa, et al.
Pubblicazione: (2024)
di: Neshaei, Seyed Parsa, et al.
Pubblicazione: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
di: Qin, Haotong, et al.
Pubblicazione: (2024)
di: Qin, Haotong, et al.
Pubblicazione: (2024)
A data science and machine learning approach to continuous analysis of Shakespeare's plays
di: Swisher, Charles, et al.
Pubblicazione: (2023)
di: Swisher, Charles, et al.
Pubblicazione: (2023)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
di: Shoham, Ofir Ben
Pubblicazione: (2026)
di: Shoham, Ofir Ben
Pubblicazione: (2026)
MenakBERT -- Hebrew Diacriticizer
di: Cohen, Ido, et al.
Pubblicazione: (2024)
di: Cohen, Ido, et al.
Pubblicazione: (2024)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
di: Deng, Jianing, et al.
Pubblicazione: (2026)
di: Deng, Jianing, et al.
Pubblicazione: (2026)
On Affine Homotopy between Language Encoders
di: Chan, Robin SM, et al.
Pubblicazione: (2024)
di: Chan, Robin SM, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Data Generation for Hardware-Friendly Post-Training Quantization
di: Dikstein, Lior, et al.
Pubblicazione: (2024) -
EPTQ: Enhanced Post-Training Quantization via Hessian-guided Network-wise Optimization
di: Gordon, Ofir, et al.
Pubblicazione: (2023) -
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
di: Gordon, Ofir, et al.
Pubblicazione: (2025) -
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
di: Achituve, Idan, et al.
Pubblicazione: (2025) -
Efficient Image Restoration via Latent Consistency Flow Matching
di: Cohen, Elad, et al.
Pubblicazione: (2025)