MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yu, Wang, Mingzi, Zou, Lancheng, Liu, Wulong, Zhen, Hui-Ling, Yuan, Mingxuan, Yu, Bei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation
by: Yin, Shuo, et al.
Published: (2026)
by: Yin, Shuo, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
by: Zhong, Ruizhe, et al.
Published: (2023)
by: Zhong, Ruizhe, et al.
Published: (2023)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
by: Wang, Kainan, et al.
Published: (2025)
by: Wang, Kainan, et al.
Published: (2025)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
by: Li, Cong, et al.
Published: (2026)
by: Li, Cong, et al.
Published: (2026)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
CPPL: A Circuit Prompt Programming Language
by: Yin, Shuo, et al.
Published: (2026)
by: Yin, Shuo, et al.
Published: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
by: Xia, Tianhua, et al.
Published: (2025)
by: Xia, Tianhua, et al.
Published: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
by: Liao, Yuchao, et al.
Published: (2026)
by: Liao, Yuchao, et al.
Published: (2026)
Hardware-Efficient FPGA Implementation of Sigmoid Function Using Mixed-Radix Hyperbolic Rotation CORDIC
by: Panchal, Chintan, et al.
Published: (2026)
by: Panchal, Chintan, et al.
Published: (2026)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
by: Errabii, Sohaib, et al.
Published: (2026)
by: Errabii, Sohaib, et al.
Published: (2026)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024)
by: Krestinskaya, Olga, et al.
Published: (2024)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024)
by: Zeng, Shulin, et al.
Published: (2024)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
by: Mamo, Oteo, et al.
Published: (2026)
by: Mamo, Oteo, et al.
Published: (2026)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
by: Wang, Peipei, et al.
Published: (2025)
by: Wang, Peipei, et al.
Published: (2025)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
by: Tsai, Pei-Huan, et al.
Published: (2026)
by: Tsai, Pei-Huan, et al.
Published: (2026)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
by: Hsieh, Fen-Yu, et al.
Published: (2025)
by: Hsieh, Fen-Yu, et al.
Published: (2025)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
by: He, Mingxuan, et al.
Published: (2024)
by: He, Mingxuan, et al.
Published: (2024)
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)
by: Wu, Hui, et al.
Published: (2023)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
by: Zhang, Yuanpeng, et al.
Published: (2025)
by: Zhang, Yuanpeng, et al.
Published: (2025)
Timing-Driven Global Placement by Efficient Critical Path Extraction
by: Shi, Yunqi, et al.
Published: (2025)
by: Shi, Yunqi, et al.
Published: (2025)
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
by: Zhang, Yipu, et al.
Published: (2025)
by: Zhang, Yipu, et al.
Published: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
Towards LLM-based Root Cause Analysis of Hardware Design Failures
by: Qiu, Siyu, et al.
Published: (2025)
by: Qiu, Siyu, et al.
Published: (2025)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
Sparsity-Aware Hardware-Software Co-Design of Spiking Neural Networks: An Overview
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
LLM-Enhanced Bayesian Optimization for Efficient Analog Layout Constraint Generation
by: Chen, Guojin, et al.
Published: (2024)
by: Chen, Guojin, et al.
Published: (2024)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
by: Kim, Dong Eun, et al.
Published: (2025)
by: Kim, Dong Eun, et al.
Published: (2025)
Similar Items
-
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025) -
PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation
by: Yin, Shuo, et al.
Published: (2026) -
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023) -
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
by: Zhong, Ruizhe, et al.
Published: (2023) -
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)