Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Janghwan, Kim, Minsoo, Baek, Seungcheol, Hwang, Seok Joong, Sung, Wonyong, Choi, Jungwook |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024)
by: Lee, Geonho, et al.
Published: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
by: Park, Seungcheol, et al.
Published: (2023)
by: Park, Seungcheol, et al.
Published: (2023)
A Comprehensive Survey of Compression Algorithms for Language Models
by: Park, Seungcheol, et al.
Published: (2024)
by: Park, Seungcheol, et al.
Published: (2024)
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
by: Trirat, Patara, et al.
Published: (2024)
by: Trirat, Patara, et al.
Published: (2024)
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
by: Baek, In-Chang, et al.
Published: (2025)
by: Baek, In-Chang, et al.
Published: (2025)
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
by: Lee, Changhun, et al.
Published: (2023)
by: Lee, Changhun, et al.
Published: (2023)
Disentangling Questions from Query Generation for Task-Adaptive Retrieval
by: Lee, Yoonsang, et al.
Published: (2024)
by: Lee, Yoonsang, et al.
Published: (2024)
Rethinking Code Refinement: Learning to Judge Code Efficiency
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
by: Jung, Yeonjoon, et al.
Published: (2024)
by: Jung, Yeonjoon, et al.
Published: (2024)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025)
by: Lee, Woongkyu, et al.
Published: (2025)
CoEx -- Co-evolving World-model and Exploration
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Efficient Real-time Refinement of Language Model Text Generation
by: Ko, Joonho, et al.
Published: (2025)
by: Ko, Joonho, et al.
Published: (2025)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
Chaining Event Spans for Temporal Relation Grounding
by: Kim, Jongho, et al.
Published: (2025)
by: Kim, Jongho, et al.
Published: (2025)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
System Prompt Optimization with Meta-Learning
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
by: Seo, Minju, et al.
Published: (2025)
by: Seo, Minju, et al.
Published: (2025)
AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
by: Lee, Sangjun, et al.
Published: (2025)
by: Lee, Sangjun, et al.
Published: (2025)
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
by: Lee, Jonghyun, et al.
Published: (2025)
by: Lee, Jonghyun, et al.
Published: (2025)
Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning
by: Kim, Jisu, et al.
Published: (2024)
by: Kim, Jisu, et al.
Published: (2024)
Low-Resource Cross-Lingual Summarization through Few-Shot Learning with Large Language Models
by: Park, Gyutae, et al.
Published: (2024)
by: Park, Gyutae, et al.
Published: (2024)
Exploring Large Language Models on Cross-Cultural Values in Connection with Training Methodology
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
Personality Editing for Language Models through Adjusting Self-Referential Queries
by: Hwang, Seojin, et al.
Published: (2025)
by: Hwang, Seojin, et al.
Published: (2025)
GWQ: Gradient-Aware Weight Quantization for Large Language Models
by: Shao, Yihua, et al.
Published: (2024)
by: Shao, Yihua, et al.
Published: (2024)
Intriguing Properties of Large Language and Vision Models
by: Lee, Young-Jun, et al.
Published: (2024)
by: Lee, Young-Jun, et al.
Published: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
by: Young, Sean I.
Published: (2024)
by: Young, Sean I.
Published: (2024)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
by: Lee, Minjae, et al.
Published: (2023)
by: Lee, Minjae, et al.
Published: (2023)
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
by: Kim, Yeeun, et al.
Published: (2024)
by: Kim, Yeeun, et al.
Published: (2024)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
by: Guo, Yipin, et al.
Published: (2024)
by: Guo, Yipin, et al.
Published: (2024)
Similar Items
-
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024) -
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024) -
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024) -
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
by: Park, Seungcheol, et al.
Published: (2025) -
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
by: Park, Seungcheol, et al.
Published: (2023)