AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Janghwan, Park, Jiwoong, Kim, Jinseok, Kim, Yongjik, Oh, Jungju, Oh, Jinwook, Choi, Jungwook |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024)
by: Lee, Geonho, et al.
Published: (2024)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
by: Lee, Minjae, et al.
Published: (2023)
by: Lee, Minjae, et al.
Published: (2023)
Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
by: Park, Jiwoong, et al.
Published: (2025)
by: Park, Jiwoong, et al.
Published: (2025)
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
by: Lee, Janghwan, et al.
Published: (2023)
by: Lee, Janghwan, et al.
Published: (2023)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025)
by: Lee, Woongkyu, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Microscaling Floating Point Formats for Large Language Models
by: Cococcioni, Marco, et al.
Published: (2025)
by: Cococcioni, Marco, et al.
Published: (2025)
Scalable Beamforming Design for Multi-RIS-Aided MU-MIMO Systems with Imperfect CSIT
by: Oh, Mintaek, et al.
Published: (2025)
by: Oh, Mintaek, et al.
Published: (2025)
Hybrid Precoding Revisited: Low-Dimensional Subspace Perspective for MU-MIMO Systems
by: Oh, Mintaek, et al.
Published: (2025)
by: Oh, Mintaek, et al.
Published: (2025)
RSCF: Relation-Semantics Consistent Filter for Entity Embedding of Knowledge Graph
by: Kim, Junsik, et al.
Published: (2025)
by: Kim, Junsik, et al.
Published: (2025)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
by: Yun, Juyoung, et al.
Published: (2023)
by: Yun, Juyoung, et al.
Published: (2023)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
by: Choi, Kanghyun, et al.
Published: (2025)
by: Choi, Kanghyun, et al.
Published: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
On the Effectiveness of Supervision in Asymmetric Non-Contrastive Learning
by: Oh, Jeongheon, et al.
Published: (2024)
by: Oh, Jeongheon, et al.
Published: (2024)
ARES: Auxiliary Range Expansion for Outlier Synthesis
by: Jung, Eui-Soo, et al.
Published: (2025)
by: Jung, Eui-Soo, et al.
Published: (2025)
Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss
by: Park, Sangyeon, et al.
Published: (2025)
by: Park, Sangyeon, et al.
Published: (2025)
Attention-aware Semantic Communications for Collaborative Inference
by: Im, Jiwoong, et al.
Published: (2024)
by: Im, Jiwoong, et al.
Published: (2024)
Full-Duplex Multiuser MISO Under Coarse Quantization: Per-Antenna SQNR Analysis and Beamforming Design
by: Yoo, Seunghyeong, et al.
Published: (2024)
by: Yoo, Seunghyeong, et al.
Published: (2024)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
by: Park, Dahoon, et al.
Published: (2026)
by: Park, Dahoon, et al.
Published: (2026)
Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility
by: Oh, Jungsuk, et al.
Published: (2026)
by: Oh, Jungsuk, et al.
Published: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Expressive Power of ReLU and Step Networks under Floating-Point Operations
by: Park, Yeachan, et al.
Published: (2024)
by: Park, Yeachan, et al.
Published: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs
by: Park, Hyobin, et al.
Published: (2026)
by: Park, Hyobin, et al.
Published: (2026)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
by: Choi, Youngwon, et al.
Published: (2026)
by: Choi, Youngwon, et al.
Published: (2026)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
A 28 nm AI microcontroller with tightly coupled zero-standby power weight memory featuring standard logic compatible 4 Mb 4-bits/cell embedded flash technology
by: Kim, Daewung, et al.
Published: (2025)
by: Kim, Daewung, et al.
Published: (2025)
SODE: Analyzing Social Dynamics in LLM Agents
by: Jung, Inseo, et al.
Published: (2026)
by: Jung, Inseo, et al.
Published: (2026)
LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
by: Lee, Yong Oh, et al.
Published: (2025)
by: Lee, Yong Oh, et al.
Published: (2025)
Towards Dynamic Trend Filtering through Trend Point Detection with Reinforcement Learning
by: Seong, Jihyeon, et al.
Published: (2024)
by: Seong, Jihyeon, et al.
Published: (2024)
Scalable and Convergent Generalized Power Iteration Precoding for Massive MIMO Systems
by: Yoo, Seunghyeong, et al.
Published: (2026)
by: Yoo, Seunghyeong, et al.
Published: (2026)
HiFloat4 Format for Language Model Inference
by: Luo, Yuanyong, et al.
Published: (2026)
by: Luo, Yuanyong, et al.
Published: (2026)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Similar Items
-
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024) -
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
by: Lee, Minjae, et al.
Published: (2023) -
Enhancing Generalization in Data-free Quantization via Mixup-class Prompting
by: Park, Jiwoong, et al.
Published: (2025) -
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
by: Lee, Janghwan, et al.
Published: (2023) -
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)