Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Heo, Jung Hwan, Kim, Jeonghoon, Kwon, Beomseok, Kim, Byeongwook, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
von: Park, Gunho, et al.
Veröffentlicht: (2022)
von: Park, Gunho, et al.
Veröffentlicht: (2022)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2024)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2024)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
von: Li, Zhikai, et al.
Veröffentlicht: (2026)
von: Li, Zhikai, et al.
Veröffentlicht: (2026)
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
von: Kim, Minsung, et al.
Veröffentlicht: (2025)
von: Kim, Minsung, et al.
Veröffentlicht: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
Label-Noise Robust Diffusion Models
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
von: Na, Byeonghu, et al.
Veröffentlicht: (2024)
GainAdaptor: Learning Quadrupedal Locomotion with Dual Actors for Adaptable and Energy-Efficient Walking on Various Terrains
von: Kim, Mincheol, et al.
Veröffentlicht: (2024)
von: Kim, Mincheol, et al.
Veröffentlicht: (2024)
Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
von: Hoang, Dung Anh, et al.
Veröffentlicht: (2025)
von: Hoang, Dung Anh, et al.
Veröffentlicht: (2025)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
von: Lee, Changhun, et al.
Veröffentlicht: (2023)
von: Lee, Changhun, et al.
Veröffentlicht: (2023)
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
von: Chung, Woojin, et al.
Veröffentlicht: (2025)
von: Chung, Woojin, et al.
Veröffentlicht: (2025)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
von: Kim, Seung-Wook, et al.
Veröffentlicht: (2025)
von: Kim, Seung-Wook, et al.
Veröffentlicht: (2025)
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
von: Wang, Jinguang, et al.
Veröffentlicht: (2025)
von: Wang, Jinguang, et al.
Veröffentlicht: (2025)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
von: Lee, Dongyeun, et al.
Veröffentlicht: (2025)
von: Lee, Dongyeun, et al.
Veröffentlicht: (2025)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
von: Lee, Geonho, et al.
Veröffentlicht: (2024)
von: Lee, Geonho, et al.
Veröffentlicht: (2024)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
Explaining How Quantization Disparately Skews a Model
von: Bellam, Abhimanyu, et al.
Veröffentlicht: (2025)
von: Bellam, Abhimanyu, et al.
Veröffentlicht: (2025)
Adversarial Bandits against Arbitrary Strategies
von: Kim, Jung-hun, et al.
Veröffentlicht: (2022)
von: Kim, Jung-hun, et al.
Veröffentlicht: (2022)
Faster Inference of LLMs using FP8 on the Intel Gaudi
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
Task Vector Quantization for Memory-Efficient Model Merging
von: Kim, Youngeun, et al.
Veröffentlicht: (2025)
von: Kim, Youngeun, et al.
Veröffentlicht: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
von: Ahn, Beomjin, et al.
Veröffentlicht: (2026)
von: Ahn, Beomjin, et al.
Veröffentlicht: (2026)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
von: Jeon, Hyesung, et al.
Veröffentlicht: (2025)
von: Jeon, Hyesung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025) -
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023) -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
von: Park, Gunho, et al.
Veröffentlicht: (2022) -
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025) -
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)