FPTQuant: Function-Preserving Transforms for LLM Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | van Breugel, Boris, Bondarenko, Yelysei, Whatmough, Paul, Nagel, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dissecting Quantization Error: A Concentration-Alignment Perspective
by: Federici, Marco, et al.
Published: (2026)
by: Federici, Marco, et al.
Published: (2026)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024)
by: van Baalen, Mart, et al.
Published: (2024)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Why Tabular Foundation Models Should Be a Research Priority
by: van Breugel, Boris, et al.
Published: (2024)
by: van Breugel, Boris, et al.
Published: (2024)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
by: Seedat, Nabeel, et al.
Published: (2023)
by: Seedat, Nabeel, et al.
Published: (2023)
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023)
by: Kuzmin, Andrey, et al.
Published: (2023)
Position: All Current Generative Fidelity and Diversity Metrics are Flawed
by: Räisä, Ossi, et al.
Published: (2025)
by: Räisä, Ossi, et al.
Published: (2025)
Efficient Reasoning on the Edge
by: Bondarenko, Yelysei, et al.
Published: (2026)
by: Bondarenko, Yelysei, et al.
Published: (2026)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
by: Skliar, Andrii, et al.
Published: (2024)
by: Skliar, Andrii, et al.
Published: (2024)
LaTable: Towards Large Tabular Models
by: van Breugel, Boris, et al.
Published: (2024)
by: van Breugel, Boris, et al.
Published: (2024)
Soft Mixture Denoising: Beyond the Expressive Bottleneck of Diffusion Models
by: Li, Yangming, et al.
Published: (2023)
by: Li, Yangming, et al.
Published: (2023)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
by: Kuzina, Anna, et al.
Published: (2025)
by: Kuzina, Anna, et al.
Published: (2025)
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
Provable Quantization with Randomized Hadamard Transform
by: Feng, Ying, et al.
Published: (2026)
by: Feng, Ying, et al.
Published: (2026)
Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
by: Mirvakhabova, Leyla, et al.
Published: (2025)
by: Mirvakhabova, Leyla, et al.
Published: (2025)
Myosotis: structured computation for attention like layer
by: Egorov, Evgenii, et al.
Published: (2025)
by: Egorov, Evgenii, et al.
Published: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
by: Cho, Yoonjun, et al.
Published: (2026)
by: Cho, Yoonjun, et al.
Published: (2026)
Privacy-Preserving Inference for Quantized BERT Models
by: Lu, Tianpei, et al.
Published: (2025)
by: Lu, Tianpei, et al.
Published: (2025)
Towards a More Complete Theory of Function Preserving Transforms
by: Painter, Michael
Published: (2024)
by: Painter, Michael
Published: (2024)
Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators
by: Tyagi, Abhishek, et al.
Published: (2022)
by: Tyagi, Abhishek, et al.
Published: (2022)
The Quantization Benefits of Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2026)
by: Ji, Yiping, et al.
Published: (2026)
Insights Into the Inner Workings of Transformer Models for Protein Function Prediction
by: Wenzel, Markus, et al.
Published: (2023)
by: Wenzel, Markus, et al.
Published: (2023)
Privacy-Preserving Quantized Federated Learning with Diverse Precision
by: Nguyen, Dang Qua, et al.
Published: (2025)
by: Nguyen, Dang Qua, et al.
Published: (2025)
AlphaJet: Automated Conceptual Aircraft Synthesis via Disentangled Generative Priors and Topology-Preserving Evolutionary Search
by: Kriuk, Boris
Published: (2026)
by: Kriuk, Boris
Published: (2026)
KurTail : Kurtosis-based LLM Quantization
by: Akhondzadeh, Mohammad Sadegh, et al.
Published: (2025)
by: Akhondzadeh, Mohammad Sadegh, et al.
Published: (2025)
Quantization-Free Autoregressive Action Transformer
by: Sheebaelhamd, Ziyad, et al.
Published: (2025)
by: Sheebaelhamd, Ziyad, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
by: Peng, Yanxin, et al.
Published: (2025)
by: Peng, Yanxin, et al.
Published: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Rethinking Residual Errors in Compensation-based LLM Quantization
by: Li, Shuaiting, et al.
Published: (2026)
by: Li, Shuaiting, et al.
Published: (2026)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
LoopQ: Quantization for Recursive Transformers
by: Fang, Rui, et al.
Published: (2026)
by: Fang, Rui, et al.
Published: (2026)
Similar Items
-
Dissecting Quantization Error: A Concentration-Alignment Perspective
by: Federici, Marco, et al.
Published: (2026) -
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025) -
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024) -
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026) -
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
by: Federici, Marco, et al.
Published: (2025)