GPTVQ: The Blessing of Dimensionality for LLM Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | van Baalen, Mart, Kuzmin, Andrey, Koryakovskiy, Ivan, Nagel, Markus, Couperus, Peter, Bastoul, Cedric, Mahurin, Eric, Blankevoort, Tijmen, Whatmough, Paul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pruning vs Quantization: Which is Better?
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
The LLM Surgeon
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2023)
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2023)
Leech Lattice Vector Quantization for Efficient LLM Compression
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2026)
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2026)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
von: Federici, Marco, et al.
Veröffentlicht: (2024)
von: Federici, Marco, et al.
Veröffentlicht: (2024)
FPTQuant: Function-Preserving Transforms for LLM Quantization
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
Dissecting Quantization Error: A Concentration-Alignment Perspective
von: Federici, Marco, et al.
Veröffentlicht: (2026)
von: Federici, Marco, et al.
Veröffentlicht: (2026)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
VeRA: Vector-based Random Matrix Adaptation
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2023)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2023)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2026)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2026)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
von: Cook, Jack, et al.
Veröffentlicht: (2025)
von: Cook, Jack, et al.
Veröffentlicht: (2025)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
SpinQuant: LLM quantization with learned rotations
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
Poro 34B and the Blessing of Multilinguality
von: Luukkonen, Risto, et al.
Veröffentlicht: (2024)
von: Luukkonen, Risto, et al.
Veröffentlicht: (2024)
Sparse High Rank Adapters
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
Clean for Haskell Programmers
von: Lubbers, Mart, et al.
Veröffentlicht: (2024)
von: Lubbers, Mart, et al.
Veröffentlicht: (2024)
Efficient Reasoning on the Edge
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2026)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2026)
Blessing of Multilinguality: A Systematic Analysis of Multilingual In-Context Learning
von: Tu, Yilei, et al.
Veröffentlicht: (2025)
von: Tu, Yilei, et al.
Veröffentlicht: (2025)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
von: Consolaro, Gianpietro, et al.
Veröffentlicht: (2024)
von: Consolaro, Gianpietro, et al.
Veröffentlicht: (2024)
Blessing or curse? A survey on the Impact of Generative AI on Fake News
von: Loth, Alexander, et al.
Veröffentlicht: (2024)
von: Loth, Alexander, et al.
Veröffentlicht: (2024)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
SqueezeLLM: Dense-and-Sparse Quantization
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
QAQ: Quality Adaptive Quantization for LLM KV Cache
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
von: Liang, Qiyao, et al.
Veröffentlicht: (2026)
von: Liang, Qiyao, et al.
Veröffentlicht: (2026)
Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models
von: Kuzmin, Gleb, et al.
Veröffentlicht: (2024)
von: Kuzmin, Gleb, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Factorio Learning Environment
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
Catastrophic Failure of LLM Unlearning via Quantization
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
The Blessing and Curse of Dimensionality in Safety Alignment
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
OP-LoRA: The Blessing of Dimensionality
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Pruning vs Quantization: Which is Better?
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023) -
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022) -
The LLM Surgeon
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2023) -
Leech Lattice Vector Quantization for Efficient LLM Compression
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2026) -
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
von: Federici, Marco, et al.
Veröffentlicht: (2024)