FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jung Hyun, Kim, Jeonghoon, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
by: Lee, Jung Hyun, et al.
Published: (2025)
by: Lee, Jung Hyun, et al.
Published: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization
by: Cha, Seohyeon, et al.
Published: (2026)
by: Cha, Seohyeon, et al.
Published: (2026)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
by: Kim, Junhan, et al.
Published: (2026)
by: Kim, Junhan, et al.
Published: (2026)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Model-Preserving Adaptive Rounding
by: Tseng, Albert, et al.
Published: (2025)
by: Tseng, Albert, et al.
Published: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
by: Gwak, Minju, et al.
Published: (2026)
by: Gwak, Minju, et al.
Published: (2026)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
by: Fang, Xingli, et al.
Published: (2026)
by: Fang, Xingli, et al.
Published: (2026)
Single-Round Scalable Analytic Federated Learning
by: Bacellar, Alan T. L., et al.
Published: (2025)
by: Bacellar, Alan T. L., et al.
Published: (2025)
Efficient Training on Multiple Consumer GPUs with RoundPipe
by: Luo, Yibin, et al.
Published: (2026)
by: Luo, Yibin, et al.
Published: (2026)
FAAR: Format-Aware Adaptive Rounding for NVFP4
by: Li, Hanglin, et al.
Published: (2026)
by: Li, Hanglin, et al.
Published: (2026)
DeepHQ: Learned Hierarchical Quantizer for Progressive Deep Image Coding
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
Broadband Ground Motion Synthesis by Diffusion Model with Minimal Condition
by: Jung, Jaeheun, et al.
Published: (2024)
by: Jung, Jaeheun, et al.
Published: (2024)
BoA: Attention-aware Post-training Quantization without Backpropagation
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
Uncertainty Calibration with Energy Based Instance-wise Scaling in the Wild Dataset
by: Kim, Mijoo, et al.
Published: (2024)
by: Kim, Mijoo, et al.
Published: (2024)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026)
by: Oh, Sehyeon, et al.
Published: (2026)
CoAst: Validation-Free Contribution Assessment for Federated Learning based on Cross-Round Valuation
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
by: Fu, Hengyu, et al.
Published: (2025)
by: Fu, Hengyu, et al.
Published: (2025)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
by: Lee, Yousung, et al.
Published: (2026)
by: Lee, Yousung, et al.
Published: (2026)
Explaining How Quantization Disparately Skews a Model
by: Bellam, Abhimanyu, et al.
Published: (2025)
by: Bellam, Abhimanyu, et al.
Published: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
by: Kwon, Sangwoo, et al.
Published: (2025)
by: Kwon, Sangwoo, et al.
Published: (2025)
Cluster-Aware Multi-Round Update for Wireless Federated Learning in Heterogeneous Environments
by: Sun, Pengcheng, et al.
Published: (2025)
by: Sun, Pengcheng, et al.
Published: (2025)
On Stochastic Rounding with Few Random Bits
by: Fitzgibbon, Andrew, et al.
Published: (2025)
by: Fitzgibbon, Andrew, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
Similar Items
-
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024) -
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024) -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025) -
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
by: Heo, Jung Hwan, et al.
Published: (2023) -
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)