AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Yanbiao, Shi, Huihong, Shao, Haikuo, Wang, Zhongfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025)
by: Ma, Shaobo, et al.
Published: (2025)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
MC$^2$A: Enabling Algorithm-Hardware Co-Design for Efficient Markov Chain Monte Carlo Acceleration
by: Zhao, Shirui, et al.
Published: (2025)
by: Zhao, Shirui, et al.
Published: (2025)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
by: Sabih, Muhammad, et al.
Published: (2025)
by: Sabih, Muhammad, et al.
Published: (2025)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
by: Liao, Yuchao, et al.
Published: (2026)
by: Liao, Yuchao, et al.
Published: (2026)
AIRCHITECT v2: Learning the Hardware Accelerator Design Space through Unified Representations
by: Seo, Jamin, et al.
Published: (2025)
by: Seo, Jamin, et al.
Published: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
by: Hwang, Ranggi, et al.
Published: (2023)
by: Hwang, Ranggi, et al.
Published: (2023)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Learning to Compare Hardware Designs for High-Level Synthesis
by: Bai, Yunsheng, et al.
Published: (2024)
by: Bai, Yunsheng, et al.
Published: (2024)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
Challenges and Research Directions for Large Language Model Inference Hardware
by: Ma, Xiaoyu, et al.
Published: (2026)
by: Ma, Xiaoyu, et al.
Published: (2026)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
by: Baldi, T., et al.
Published: (2026)
by: Baldi, T., et al.
Published: (2026)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Architect in the Loop Agentic Hardware Design and Verification
by: Mohammed, Mubarek
Published: (2025)
by: Mohammed, Mubarek
Published: (2025)
ML For Hardware Design Interpretability: Challenges and Opportunities
by: Baartmans, Raymond, et al.
Published: (2025)
by: Baartmans, Raymond, et al.
Published: (2025)
HDReason: Algorithm-Hardware Codesign for Hyperdimensional Knowledge Graph Reasoning
by: Chen, Hanning, et al.
Published: (2024)
by: Chen, Hanning, et al.
Published: (2024)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
by: Xiang, Maoyang, et al.
Published: (2025)
by: Xiang, Maoyang, et al.
Published: (2025)
HiVeGen -- Hierarchical LLM-based Verilog Generation for Scalable Chip Design
by: Tang, Jinwei, et al.
Published: (2024)
by: Tang, Jinwei, et al.
Published: (2024)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
by: Zhang, Yuanpeng, et al.
Published: (2025)
by: Zhang, Yuanpeng, et al.
Published: (2025)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
by: Ahmed, Md Rubel, et al.
Published: (2024)
by: Ahmed, Md Rubel, et al.
Published: (2024)
PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection
by: Park, Hyoseok, et al.
Published: (2026)
by: Park, Hyoseok, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024)
by: Krestinskaya, Olga, et al.
Published: (2024)
Similar Items
-
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025) -
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024) -
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024) -
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024) -
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)