Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nrusimha, Aniruddha, Mishra, Mayank, Wang, Naigang, Alistarh, Dan, Panda, Rameswar, Kim, Yoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
Towards Verifiable Text Generation with Symbolic References
von: Hennigen, Lucas Torroba, et al.
Veröffentlicht: (2023)
von: Hennigen, Lucas Torroba, et al.
Veröffentlicht: (2023)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
Compression Scaling Laws:Unifying Sparsity and Quantization
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
ECO: Quantized Training without Full-Precision Master Weights
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning
von: Guo, Han, et al.
Veröffentlicht: (2023)
von: Guo, Han, et al.
Veröffentlicht: (2023)
Beyond Outliers: A Study of Optimizers Under Quantization
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Systematic Outliers in Large Language Models
von: An, Yongqi, et al.
Veröffentlicht: (2025)
von: An, Yongqi, et al.
Veröffentlicht: (2025)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
Frayed RoPE and Long Inputs: A Geometric Perspective
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
von: Paglieri, Davide, et al.
Veröffentlicht: (2024)
von: Paglieri, Davide, et al.
Veröffentlicht: (2024)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
von: Jeon, Hyesung, et al.
Veröffentlicht: (2024)
von: Jeon, Hyesung, et al.
Veröffentlicht: (2024)
Calibrating Expressions of Certainty
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
von: van der Wal, Oskar, et al.
Veröffentlicht: (2025)
von: van der Wal, Oskar, et al.
Veröffentlicht: (2025)
MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores
von: Zhou, Fengwei, et al.
Veröffentlicht: (2025)
von: Zhou, Fengwei, et al.
Veröffentlicht: (2025)
Learning to Decode Collaboratively with Multiple Language Models
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2024)
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2024)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024) -
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025) -
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024) -
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025) -
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)