GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jinuk, Halabi, Marwa El, Park, Wonpyo, Schaefer, Clemens JS, Lee, Deokjae, Park, Yeonhong, Lee, Jae W., Song, Hyun Oh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
LayerMerge: Neural Network Depth Compression through Layer Pruning and Merging
by: Kim, Jinuk, et al.
Published: (2024)
by: Kim, Jinuk, et al.
Published: (2024)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024)
by: Lee, Deokjae, et al.
Published: (2024)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
by: Lee, Haeun, et al.
Published: (2025)
by: Lee, Haeun, et al.
Published: (2025)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
by: Shin, Sungbin, et al.
Published: (2024)
by: Shin, Sungbin, et al.
Published: (2024)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
by: Park, Bumsu, et al.
Published: (2026)
by: Park, Bumsu, et al.
Published: (2026)
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation
by: Kim, Jinuk, et al.
Published: (2026)
by: Kim, Jinuk, et al.
Published: (2026)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
by: Kwon, Sangwoo, et al.
Published: (2025)
by: Kwon, Sangwoo, et al.
Published: (2025)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
3DTurboQuant: Training-Free Near-Optimal Quantization for 3D Reconstruction Models
by: Lee, Jae Joong
Published: (2026)
by: Lee, Jae Joong
Published: (2026)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
ESC-MVQ: End-to-End Semantic Communication With Multi-Codebook Vector Quantization
by: Shin, Junyong, et al.
Published: (2025)
by: Shin, Junyong, et al.
Published: (2025)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
by: Baras, Amit, et al.
Published: (2023)
by: Baras, Amit, et al.
Published: (2023)
Difference of Submodular Minimization via DC Programming
by: Halabi, Marwa El, et al.
Published: (2023)
by: Halabi, Marwa El, et al.
Published: (2023)
Discrete and Continuous Difference of Submodular Minimization
by: Orfanides, George, et al.
Published: (2025)
by: Orfanides, George, et al.
Published: (2025)
Learning to Better Search with Language Models via Guided Reinforced Self-Training
by: Moon, Seungyong, et al.
Published: (2024)
by: Moon, Seungyong, et al.
Published: (2024)
HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs
by: An, Sumin, et al.
Published: (2025)
by: An, Sumin, et al.
Published: (2025)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Exploiting Fine-Grained Skip Behaviors for Micro-Video Recommendation
by: Lee, Sanghyuck, et al.
Published: (2025)
by: Lee, Sanghyuck, et al.
Published: (2025)
MVAdapt: Zero-Shot Multi-Vehicle Adaptation for End-to-End Autonomous Driving
by: Oh, Haesung, et al.
Published: (2026)
by: Oh, Haesung, et al.
Published: (2026)
Conversational Query Reformulation with the Guidance of Retrieved Documents
by: Park, Jeonghyun, et al.
Published: (2024)
by: Park, Jeonghyun, et al.
Published: (2024)
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
by: Lee, Jemin, et al.
Published: (2024)
by: Lee, Jemin, et al.
Published: (2024)
Efficacy of plant‐derived dietary supplements in improving overall menopausal symptoms in women: An updated systematic review and meta‐analysis
by: Mi Ra Oh, et al.
Published: (2024)
by: Mi Ra Oh, et al.
Published: (2024)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
CAT: Contrastive Adapter Training for Personalized Image Generation
by: Park, Jae Wan, et al.
Published: (2024)
by: Park, Jae Wan, et al.
Published: (2024)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
by: Lee, Jonghun, et al.
Published: (2026)
by: Lee, Jonghun, et al.
Published: (2026)
Underactuated Robotic Hand with Grasp State Estimation Using Tendon-Based Proprioception
by: Lee, Jae-Hyun, et al.
Published: (2025)
by: Lee, Jae-Hyun, et al.
Published: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
Clinical Effectiveness of Dupilumab in Eosinophilic Granulomatosis With Polyangiitis: A Retrospective Observational Study
by: Oh Chan Kwon, et al.
Published: (2026)
by: Oh Chan Kwon, et al.
Published: (2026)
Efficient One‐Step Production of 7S,17S‐ and 10S,17S‐Dihydroxydocosahexaenoic Acids by a Double‐Oxygenating 15S‐Lipoxygenase From Chlamydomonas incerta
by: Hyun‐Ah Park, et al.
Published: (2025)
by: Hyun‐Ah Park, et al.
Published: (2025)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
Chunk-Guided Q-Learning
by: Song, Gwanwoo, et al.
Published: (2026)
by: Song, Gwanwoo, et al.
Published: (2026)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
Similar Items
-
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025) -
LayerMerge: Neural Network Depth Compression through Layer Pruning and Merging
by: Kim, Jinuk, et al.
Published: (2024) -
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
by: Park, Yeonhong, et al.
Published: (2024) -
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024) -
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)