Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jinhao, Xu, Jiaming, Huang, Shan, Chen, Yonghua, Li, Wen, Liu, Jun, Lian, Yaoxiu, Pan, Jiayi, Ding, Li, Zhou, Hao, Wang, Yu, Dai, Guohao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MARCA: Mamba Accelerator with ReConfigurable Architecture
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
di: Zhang, Jinsong, et al.
Pubblicazione: (2025)
di: Zhang, Jinsong, et al.
Pubblicazione: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2024)
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2024)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
di: Lu, Jinming, et al.
Pubblicazione: (2025)
di: Lu, Jinming, et al.
Pubblicazione: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
di: Hwang, Soojin, et al.
Pubblicazione: (2025)
di: Hwang, Soojin, et al.
Pubblicazione: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
di: Li, Zhengke, et al.
Pubblicazione: (2025)
di: Li, Zhengke, et al.
Pubblicazione: (2025)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
di: Kim, Dong Eun, et al.
Pubblicazione: (2025)
di: Kim, Dong Eun, et al.
Pubblicazione: (2025)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
di: Huang, Mingqiang, et al.
Pubblicazione: (2024)
di: Huang, Mingqiang, et al.
Pubblicazione: (2024)
SimulatorCoder: DNN Accelerator Simulator Code Generation and Optimization via Large Language Models
di: Xia, Yuhuan, et al.
Pubblicazione: (2026)
di: Xia, Yuhuan, et al.
Pubblicazione: (2026)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
di: Ding, Hong, et al.
Pubblicazione: (2025)
di: Ding, Hong, et al.
Pubblicazione: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
di: Geens, Robin, et al.
Pubblicazione: (2026)
di: Geens, Robin, et al.
Pubblicazione: (2026)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
di: Zhong, Yang, et al.
Pubblicazione: (2025)
di: Zhong, Yang, et al.
Pubblicazione: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
di: Wang, Peipei, et al.
Pubblicazione: (2025)
di: Wang, Peipei, et al.
Pubblicazione: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
di: Li, Jindong, et al.
Pubblicazione: (2025)
di: Li, Jindong, et al.
Pubblicazione: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021)
di: Wang, Zhao, et al.
Pubblicazione: (2021)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
di: Li, Cong, et al.
Pubblicazione: (2026)
di: Li, Cong, et al.
Pubblicazione: (2026)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
di: Li, Boyu, et al.
Pubblicazione: (2025)
di: Li, Boyu, et al.
Pubblicazione: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
di: Guo, Cong, et al.
Pubblicazione: (2025)
di: Guo, Cong, et al.
Pubblicazione: (2025)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
di: Shahawy, Mohamed, et al.
Pubblicazione: (2025)
di: Shahawy, Mohamed, et al.
Pubblicazione: (2025)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
di: Du, Shuting, et al.
Pubblicazione: (2025)
di: Du, Shuting, et al.
Pubblicazione: (2025)
CAMASim: A Comprehensive Simulation Framework for Content-Addressable Memory based Accelerators
di: Li, Mengyuan, et al.
Pubblicazione: (2024)
di: Li, Mengyuan, et al.
Pubblicazione: (2024)
Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
di: Pan, Tianwei, et al.
Pubblicazione: (2025)
di: Pan, Tianwei, et al.
Pubblicazione: (2025)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
di: Guan, Yue, et al.
Pubblicazione: (2026)
di: Guan, Yue, et al.
Pubblicazione: (2026)
TAMI-MPC:Trusted Acceleration of Minimal-Interaction MPC for Efficient Nonlinear Inference
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators
di: Agarwal, Ayushi, et al.
Pubblicazione: (2026)
di: Agarwal, Ayushi, et al.
Pubblicazione: (2026)
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
di: Ando, Takuto, et al.
Pubblicazione: (2025)
di: Ando, Takuto, et al.
Pubblicazione: (2025)
Hardware Acceleration in Portable MRIs: State of the Art and Future Prospects
di: Habsi, Omar Al, et al.
Pubblicazione: (2025)
di: Habsi, Omar Al, et al.
Pubblicazione: (2025)
Xpikeformer: Hybrid Analog-Digital Hardware Acceleration for Spiking Transformers
di: Song, Zihang, et al.
Pubblicazione: (2024)
di: Song, Zihang, et al.
Pubblicazione: (2024)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
di: Hao, Mingbo, et al.
Pubblicazione: (2026)
di: Hao, Mingbo, et al.
Pubblicazione: (2026)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2025)
di: Lu, Tsung-Han, et al.
Pubblicazione: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
di: Chen, Jinwu, et al.
Pubblicazione: (2026)
di: Chen, Jinwu, et al.
Pubblicazione: (2026)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
di: Xuan, Zihao, et al.
Pubblicazione: (2026)
di: Xuan, Zihao, et al.
Pubblicazione: (2026)
Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints
di: Gao, Xiangyu, et al.
Pubblicazione: (2026)
di: Gao, Xiangyu, et al.
Pubblicazione: (2026)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
di: Wang, Kainan, et al.
Pubblicazione: (2025)
di: Wang, Kainan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MARCA: Mamba Accelerator with ReConfigurable Architecture
di: Li, Jinhao, et al.
Pubblicazione: (2024) -
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
di: Rakka, Mariam, et al.
Pubblicazione: (2024) -
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
di: Zhang, Jinsong, et al.
Pubblicazione: (2025) -
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2024) -
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
di: Lu, Jinming, et al.
Pubblicazione: (2025)