Improving Quantization with Post-Training Model Expansion
Fuente:
arXiv
Saved in:
| Main Authors: | Franco, Giuseppe, Monteagudo-Lago, Pablo, Colbert, Ian, Fraser, Nicholas, Blott, Michaela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
Pushing the Limits of Block Rotations in Post-Training Quantization
by: Sanjeet, Sai, et al.
Published: (2026)
by: Sanjeet, Sai, et al.
Published: (2026)
FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
by: Khandelwal, Shashwat, et al.
Published: (2025)
by: Khandelwal, Shashwat, et al.
Published: (2025)
A2Q+: Improving Accumulator-Aware Weight Quantization
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
by: Qin, Yifan, et al.
Published: (2023)
by: Qin, Yifan, et al.
Published: (2023)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
Late Breaking Results: Quamba-SE: Soft-edge Quantizer for Activations in State Space Models
by: Chen, Yizhi, et al.
Published: (2026)
by: Chen, Yizhi, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks
by: Unnikrishnan, Nanda K., et al.
Published: (2025)
by: Unnikrishnan, Nanda K., et al.
Published: (2025)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
by: Matsushima, Kosuke, et al.
Published: (2026)
by: Matsushima, Kosuke, et al.
Published: (2026)
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023)
by: Bang, Jehyeon, et al.
Published: (2023)
Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
by: Abushahla, Hamza A., et al.
Published: (2025)
by: Abushahla, Hamza A., et al.
Published: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
RESQ: A Unified Framework for REliability- and Security Enhancement of Quantized Deep Neural Networks
by: Mohammadi, Ali Soltan, et al.
Published: (2026)
by: Mohammadi, Ali Soltan, et al.
Published: (2026)
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
by: Gao, Mingzhe, et al.
Published: (2024)
by: Gao, Mingzhe, et al.
Published: (2024)
Heterogeneous Acceleration Pipeline for Recommendation System Training
by: Adnan, Muhammad, et al.
Published: (2022)
by: Adnan, Muhammad, et al.
Published: (2022)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
by: Demirkiran, Cansu, et al.
Published: (2023)
by: Demirkiran, Cansu, et al.
Published: (2023)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
by: Meng, Chang, et al.
Published: (2026)
by: Meng, Chang, et al.
Published: (2026)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
by: Killian, Earl
Published: (2026)
by: Killian, Earl
Published: (2026)
Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
by: Brandoit, Julien, et al.
Published: (2026)
by: Brandoit, Julien, et al.
Published: (2026)
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities
by: Read, James, et al.
Published: (2025)
by: Read, James, et al.
Published: (2025)
Characterizing and Understanding HGNN Training on GPUs
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
Chip Placement with Diffusion Models
by: Lee, Vint, et al.
Published: (2024)
by: Lee, Vint, et al.
Published: (2024)
HiFloat4 Format for Language Model Inference
by: Luo, Yuanyong, et al.
Published: (2026)
by: Luo, Yuanyong, et al.
Published: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
Challenges and Research Directions for Large Language Model Inference Hardware
by: Ma, Xiaoyu, et al.
Published: (2026)
by: Ma, Xiaoyu, et al.
Published: (2026)
EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion Models
by: Heo, Jaehoon, et al.
Published: (2025)
by: Heo, Jaehoon, et al.
Published: (2025)
Reducing the Barriers to Entry for Foundation Model Training
by: Faraboschi, Paolo, et al.
Published: (2024)
by: Faraboschi, Paolo, et al.
Published: (2024)
LaMAGIC: Language-Model-based Topology Generation for Analog Integrated Circuits
by: Chang, Chen-Chia, et al.
Published: (2024)
by: Chang, Chen-Chia, et al.
Published: (2024)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
by: Kim, Taehyun, et al.
Published: (2024)
by: Kim, Taehyun, et al.
Published: (2024)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
by: Yadav, Divakar Kumar, et al.
Published: (2026)
by: Yadav, Divakar Kumar, et al.
Published: (2026)
Similar Items
-
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026) -
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023) -
Pushing the Limits of Block Rotations in Post-Training Quantization
by: Sanjeet, Sai, et al.
Published: (2026) -
FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
by: Khandelwal, Shashwat, et al.
Published: (2025) -
A2Q+: Improving Accumulator-Aware Weight Quantization
by: Colbert, Ian, et al.
Published: (2024)