The LLM Surgeon
Fuente:
arXiv
Saved in:
| Main Authors: | van der Ouderaa, Tycho F. A., Nagel, Markus, van Baalen, Mart, Asano, Yuki M., Blankevoort, Tijmen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024)
by: van Baalen, Mart, et al.
Published: (2024)
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023)
by: Kuzmin, Andrey, et al.
Published: (2023)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
VeRA: Vector-based Random Matrix Adaptation
by: Kopiczko, Dawid J., et al.
Published: (2023)
by: Kopiczko, Dawid J., et al.
Published: (2023)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
by: Kopiczko, Dawid J., et al.
Published: (2024)
by: Kopiczko, Dawid J., et al.
Published: (2024)
Noether's razor: Learning Conserved Quantities
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
by: Kopiczko, Dawid J., et al.
Published: (2026)
by: Kopiczko, Dawid J., et al.
Published: (2026)
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024)
by: Gelberg, Yoav, et al.
Published: (2024)
Pyramid Vector Quantization for LLMs
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
by: Cook, Jack, et al.
Published: (2025)
by: Cook, Jack, et al.
Published: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
by: Skliar, Andrii, et al.
Published: (2024)
by: Skliar, Andrii, et al.
Published: (2024)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
Visualizing token importance for black-box language models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
Factorio Learning Environment
by: Hopkins, Jack, et al.
Published: (2025)
by: Hopkins, Jack, et al.
Published: (2025)
Language Bottleneck Models for Qualitative Knowledge State Modeling
by: Berthon, Antonin, et al.
Published: (2025)
by: Berthon, Antonin, et al.
Published: (2025)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
Agentic Large Language Models, a survey
by: Plaat, Aske, et al.
Published: (2025)
by: Plaat, Aske, et al.
Published: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
Quantifying perturbation impacts for large language models
by: Rauba, Paulius, et al.
Published: (2024)
by: Rauba, Paulius, et al.
Published: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
FPTQuant: Function-Preserving Transforms for LLM Quantization
by: van Breugel, Boris, et al.
Published: (2025)
by: van Breugel, Boris, et al.
Published: (2025)
Beyond Correlation: Refutation-Validated Aspect-Based Sentiment Analysis for Explainable Energy Market Returns
by: van der Heever, Wihan, et al.
Published: (2026)
by: van der Heever, Wihan, et al.
Published: (2026)
Active Task Disambiguation with LLMs
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
by: van Buuren, Stef
Published: (2026)
by: van Buuren, Stef
Published: (2026)
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
by: van der Wal, Oskar, et al.
Published: (2025)
by: van der Wal, Oskar, et al.
Published: (2025)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
A comparison of latent semantic analysis and correspondence analysis of document-term matrices
by: Qi, Qianqian, et al.
Published: (2021)
by: Qi, Qianqian, et al.
Published: (2021)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
by: Miyano, Ryota, et al.
Published: (2025)
by: Miyano, Ryota, et al.
Published: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
by: Schouten, Stefan F., et al.
Published: (2025)
by: Schouten, Stefan F., et al.
Published: (2025)
Similar Items
-
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026) -
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024) -
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023) -
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022) -
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)