Accurate Block Quantization in LLMs with Outliers
Fuente:
arXiv
Saved in:
| Main Authors: | Trukhanov, Nikita, Soloveychik, Ilya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Inference Acceleration via Efficient Operation Fusion
by: Salmani, Mahsa, et al.
Published: (2025)
by: Salmani, Mahsa, et al.
Published: (2025)
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
by: Agrawal, Aditya, et al.
Published: (2024)
by: Agrawal, Aditya, et al.
Published: (2024)
Accurate Models of NVIDIA Tensor Cores
by: Khattak, Faizan A., et al.
Published: (2025)
by: Khattak, Faizan A., et al.
Published: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
Design and accuracy trade-offs in Computational Statistics
by: Xu, Tiancheng, et al.
Published: (2025)
by: Xu, Tiancheng, et al.
Published: (2025)
Efficient FRW Transitions via Stochastic Finite Differences for Handling Non-Stratified Dielectrics
by: Huang, Jiechen, et al.
Published: (2025)
by: Huang, Jiechen, et al.
Published: (2025)
An Open-Source Framework for Efficient Numerically-Tailored Computations
by: Ledoux, Louis, et al.
Published: (2024)
by: Ledoux, Louis, et al.
Published: (2024)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
ScaleRTL: Scaling LLMs with Reasoning Data and Test-Time Compute for Accurate RTL Code Generation
by: Deng, Chenhui, et al.
Published: (2025)
by: Deng, Chenhui, et al.
Published: (2025)
MATLAB Simulator of Level-Index Arithmetic
by: Mikaitis, Mantas
Published: (2024)
by: Mikaitis, Mantas
Published: (2024)
Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
by: Croci, M., et al.
Published: (2024)
by: Croci, M., et al.
Published: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
PAI: Fast, Accurate, and Full Benchmark Performance Projection with AI
by: Johnson, Avery, et al.
Published: (2026)
by: Johnson, Avery, et al.
Published: (2026)
EULER-ADAS: Energy-Efficient & SIMD-Unified Logarithmic-Posit Engine for Precision-Reconfigurable Approximate ADAS Acceleration
by: Lokhande, Mukul, et al.
Published: (2026)
by: Lokhande, Mukul, et al.
Published: (2026)
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
by: Qin, Yifan, et al.
Published: (2024)
by: Qin, Yifan, et al.
Published: (2024)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
by: Matsushima, Kosuke, et al.
Published: (2026)
by: Matsushima, Kosuke, et al.
Published: (2026)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
by: Errabii, Sohaib, et al.
Published: (2026)
by: Errabii, Sohaib, et al.
Published: (2026)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
by: Kam, Dongyun, et al.
Published: (2024)
by: Kam, Dongyun, et al.
Published: (2024)
Are LLMs Any Good for High-Level Synthesis?
by: Liao, Yuchao, et al.
Published: (2024)
by: Liao, Yuchao, et al.
Published: (2024)
Hardware Acceleration of LLMs: A comprehensive survey and comparison
by: Koilia, Nikoletta, et al.
Published: (2024)
by: Koilia, Nikoletta, et al.
Published: (2024)
Classification-Based Automatic HDL Code Generation Using LLMs
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
HAVEN: Hybrid Automated Verification ENgine for UVM Testbench Synthesis with LLMs
by: Meng, Chang-Chih, et al.
Published: (2026)
by: Meng, Chang-Chih, et al.
Published: (2026)
AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs
by: Gao, Mingzhe, et al.
Published: (2024)
by: Gao, Mingzhe, et al.
Published: (2024)
iDSE: Navigating Design Space Exploration in High-Level Synthesis Using LLMs
by: Li, Runkai, et al.
Published: (2025)
by: Li, Runkai, et al.
Published: (2025)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
ChatSVA: Bridging SVA Generation for Hardware Verification via Task-Specific LLMs
by: Fu, Lik Tung, et al.
Published: (2026)
by: Fu, Lik Tung, et al.
Published: (2026)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
by: Abi-Karam, Stefan, et al.
Published: (2025)
by: Abi-Karam, Stefan, et al.
Published: (2025)
Configuration Over Selection: Hyperparameter Sensitivity Exceeds Model Differences in Open-Source LLMs for RTL Generation
by: Shao, Minghao, et al.
Published: (2026)
by: Shao, Minghao, et al.
Published: (2026)
Automatic High-quality Verilog Assertion Generation through Subtask-Focused Fine-Tuned LLMs and Iterative Prompting
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
by: Chen, Yao, et al.
Published: (2022)
by: Chen, Yao, et al.
Published: (2022)
A low-rank balanced truncation approach for large-scale RLCk model order reduction based on extended Krylov subspace and a frequency-aware convergence criterion
by: Giamouzis, Christos, et al.
Published: (2024)
by: Giamouzis, Christos, et al.
Published: (2024)
Hawkeye: Reproducing GPU-Level Non-Determinism
by: Badash, Erez, et al.
Published: (2026)
by: Badash, Erez, et al.
Published: (2026)
MORCIC: Model Order Reduction Techniques for Electromagnetic Models of Integrated Circuits
by: Garyfallou, Dimitrios, et al.
Published: (2023)
by: Garyfallou, Dimitrios, et al.
Published: (2023)
Improving Quantization with Post-Training Model Expansion
by: Franco, Giuseppe, et al.
Published: (2025)
by: Franco, Giuseppe, et al.
Published: (2025)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Similar Items
-
LLM Inference Acceleration via Efficient Operation Fusion
by: Salmani, Mahsa, et al.
Published: (2025) -
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
by: Agrawal, Aditya, et al.
Published: (2024) -
Accurate Models of NVIDIA Tensor Cores
by: Khattak, Faizan A., et al.
Published: (2025) -
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
by: Xie, Peichen, et al.
Published: (2025) -
Design and accuracy trade-offs in Computational Statistics
by: Xu, Tiancheng, et al.
Published: (2025)