EXAQ: Exponent Aware Quantization For LLMs Acceleration
Fuente:
arXiv
Saved in:
| Main Authors: | Shkolnik, Moran, Fishman, Maxim, Chmiel, Brian, Ben-Yaacov, Hilla, Banner, Ron, Levy, Kfir Yehuda |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
LLMs for Analog Circuit Design Continuum (ACDC)
by: Esfandiari, Yasaman, et al.
Published: (2025)
by: Esfandiari, Yasaman, et al.
Published: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
by: Hourri, Younes, et al.
Published: (2025)
by: Hourri, Younes, et al.
Published: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
by: Tran, Nguyen Phuc, et al.
Published: (2025)
by: Tran, Nguyen Phuc, et al.
Published: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
by: Mozaffari, Mohammad, et al.
Published: (2025)
by: Mozaffari, Mohammad, et al.
Published: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
Accelerating AI Performance using Anderson Extrapolation on GPUs
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
by: Du, Dayou, et al.
Published: (2024)
by: Du, Dayou, et al.
Published: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
by: Çöplü, Tolga, et al.
Published: (2023)
by: Çöplü, Tolga, et al.
Published: (2023)
Unleashing Automated Congestion Control Customization in the Wild
by: Cohen, Amit, et al.
Published: (2025)
by: Cohen, Amit, et al.
Published: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2023)
by: Rajput, Saurabhsingh, et al.
Published: (2023)
Kevin: Multi-Turn RL for Generating CUDA Kernels
by: Baronio, Carlo, et al.
Published: (2025)
by: Baronio, Carlo, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
by: Hangun, Batuhan, et al.
Published: (2025)
by: Hangun, Batuhan, et al.
Published: (2025)
Deploying Open-Source Large Language Models: A performance Analysis
by: Bendi-Ouis, Yannis, et al.
Published: (2024)
by: Bendi-Ouis, Yannis, et al.
Published: (2024)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
by: Jiang, Jevin, et al.
Published: (2026)
by: Jiang, Jevin, et al.
Published: (2026)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
Similar Items
-
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025) -
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024) -
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026) -
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)