LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Jung Hyun, Kim, Jeonghoon, Yang, June Yong, Kwon, Se Jung, Yang, Eunho, Yoo, Kang Min, Lee, Dongsoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
di: Lee, Jung Hyun, et al.
Pubblicazione: (2026)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
di: Yang, June Yong, et al.
Pubblicazione: (2024)
di: Yang, June Yong, et al.
Pubblicazione: (2024)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
di: Yang, Lianwei, et al.
Pubblicazione: (2025)
di: Yang, Lianwei, et al.
Pubblicazione: (2025)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
di: Lee, Jeongtae, et al.
Pubblicazione: (2026)
di: Lee, Jeongtae, et al.
Pubblicazione: (2026)
An Inquiry into Datacenter TCO for LLM Inference with FP8
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation
di: Lee, Byung Hyun, et al.
Pubblicazione: (2025)
di: Lee, Byung Hyun, et al.
Pubblicazione: (2025)
Unleashing the Potential of Text-attributed Graphs: Automatic Relation Decomposition via Large Language Models
di: Seo, Hyunjin, et al.
Pubblicazione: (2024)
di: Seo, Hyunjin, et al.
Pubblicazione: (2024)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
di: Kim, Yoonjeon, et al.
Pubblicazione: (2024)
di: Kim, Yoonjeon, et al.
Pubblicazione: (2024)
Toward Robust LiDAR based 3D Object Detection via Density-Aware Adaptive Thresholding
di: Lee, Eunho, et al.
Pubblicazione: (2024)
di: Lee, Eunho, et al.
Pubblicazione: (2024)
B2F: End-to-End Body-to-Face Motion Generation with Style Reference
di: Jang, Bokyung, et al.
Pubblicazione: (2025)
di: Jang, Bokyung, et al.
Pubblicazione: (2025)
PhysicsFC: Learning User-Controlled Skills for a Physics-Based Football Player Controller
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
Provable Low-Rank Tensor-Train Approximations in the Inverse of Large-Scale Structured Matrices
di: Xiao, Chuanfu, et al.
Pubblicazione: (2025)
di: Xiao, Chuanfu, et al.
Pubblicazione: (2025)
A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample Perspective
di: Jung, Yeonsung, et al.
Pubblicazione: (2024)
di: Jung, Yeonsung, et al.
Pubblicazione: (2024)
Label-Noise Robust Diffusion Models
di: Na, Byeonghu, et al.
Pubblicazione: (2024)
di: Na, Byeonghu, et al.
Pubblicazione: (2024)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
di: Park, Cheonbok, et al.
Pubblicazione: (2025)
di: Park, Cheonbok, et al.
Pubblicazione: (2025)
RefQSR: Reference-based Quantization for Image Super-Resolution Networks
di: Lee, Hongjae, et al.
Pubblicazione: (2024)
di: Lee, Hongjae, et al.
Pubblicazione: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
Novel methods and approaches for supercapacitor separator using waste newspaper
di: Min Jun Lee, et al.
Pubblicazione: (2025)
di: Min Jun Lee, et al.
Pubblicazione: (2025)
Monthly means of triple isotopic compositions of precipitation in Seoul (2016-2020)
di: Kim, Songyi, et al.
Pubblicazione: (2025)
di: Kim, Songyi, et al.
Pubblicazione: (2025)
Rank-O-ToM: Unlocking Emotional Nuance Ranking to Enhance Affective Theory-of-Mind
di: Kim, JiHyun, et al.
Pubblicazione: (2025)
di: Kim, JiHyun, et al.
Pubblicazione: (2025)
Divide and Translate: Compositional First-Order Logic Translation and Verification for Complex Logical Reasoning
di: Ryu, Hyun, et al.
Pubblicazione: (2024)
di: Ryu, Hyun, et al.
Pubblicazione: (2024)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
di: Liu, Yifei, et al.
Pubblicazione: (2024)
di: Liu, Yifei, et al.
Pubblicazione: (2024)
Prediction of Permissioned Blockchain Performance for Resource Scaling Configurations
di: Jung, Seungwoo, et al.
Pubblicazione: (2025)
di: Jung, Seungwoo, et al.
Pubblicazione: (2025)
Post-Training Quantization for Vision Mamba with k-Scaled Quantization and Reparameterization
di: Shi, Bo-Yun, et al.
Pubblicazione: (2025)
di: Shi, Bo-Yun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023) -
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023) -
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
di: Lee, Jung Hyun, et al.
Pubblicazione: (2026) -
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
di: Yang, June Yong, et al.
Pubblicazione: (2024) -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)