Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jindong, Li, Tenglong, Shen, Guobin, Zhao, Dongcheng, Zhang, Qian, Zeng, Yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
di: Li, Jindong, et al.
Pubblicazione: (2025)
di: Li, Jindong, et al.
Pubblicazione: (2025)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
di: Li, Jindong, et al.
Pubblicazione: (2024)
di: Li, Jindong, et al.
Pubblicazione: (2024)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
di: Li, Tenglong, et al.
Pubblicazione: (2026)
di: Li, Tenglong, et al.
Pubblicazione: (2026)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
di: Li, Tenglong, et al.
Pubblicazione: (2024)
di: Li, Tenglong, et al.
Pubblicazione: (2024)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
di: Li, Tenglong, et al.
Pubblicazione: (2025)
di: Li, Tenglong, et al.
Pubblicazione: (2025)
Pushing the Memory Bandwidth Wall with CXL-enabled Idle I/O Bandwidth Harvesting
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
di: Ferry, Corentin, et al.
Pubblicazione: (2024)
di: Ferry, Corentin, et al.
Pubblicazione: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
di: Cheng, Feng, et al.
Pubblicazione: (2025)
di: Cheng, Feng, et al.
Pubblicazione: (2025)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
di: Huang, Mingqiang, et al.
Pubblicazione: (2024)
di: Huang, Mingqiang, et al.
Pubblicazione: (2024)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
di: Yin, Chenyang, et al.
Pubblicazione: (2025)
di: Yin, Chenyang, et al.
Pubblicazione: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
FPGA-based Hyrbid Memory Emulation System
di: Wen, Fei, et al.
Pubblicazione: (2020)
di: Wen, Fei, et al.
Pubblicazione: (2020)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
di: Yu, Feng, et al.
Pubblicazione: (2026)
di: Yu, Feng, et al.
Pubblicazione: (2026)
EMiX: Emulating Beyond Single-FPGA Limits
di: Kropotov, Alexander, et al.
Pubblicazione: (2026)
di: Kropotov, Alexander, et al.
Pubblicazione: (2026)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
di: Ai, Chenyang, et al.
Pubblicazione: (2026)
di: Ai, Chenyang, et al.
Pubblicazione: (2026)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
di: Shin, Duckgyu, et al.
Pubblicazione: (2026)
di: Shin, Duckgyu, et al.
Pubblicazione: (2026)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
di: Lin, Ye, et al.
Pubblicazione: (2026)
di: Lin, Ye, et al.
Pubblicazione: (2026)
Scaling Routers with In-Package Optics and High-Bandwidth Memories
di: Keslassy, Isaac, et al.
Pubblicazione: (2026)
di: Keslassy, Isaac, et al.
Pubblicazione: (2026)
Design and Implementation of BNN-Based Object Detection on FPGA
di: Zhao, Xuyu, et al.
Pubblicazione: (2026)
di: Zhao, Xuyu, et al.
Pubblicazione: (2026)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
di: Gupta, Neelesh, et al.
Pubblicazione: (2026)
di: Gupta, Neelesh, et al.
Pubblicazione: (2026)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
di: Zhang, Weichuang, et al.
Pubblicazione: (2024)
di: Zhang, Weichuang, et al.
Pubblicazione: (2024)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
di: Wang, Ruibao, et al.
Pubblicazione: (2024)
di: Wang, Ruibao, et al.
Pubblicazione: (2024)
Per-Bank Memory Bandwidth Regulation for Predictable and Performant Real-Time System
di: Sullivan, Connor Rudy, et al.
Pubblicazione: (2026)
di: Sullivan, Connor Rudy, et al.
Pubblicazione: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
di: Atmer, Hannah, et al.
Pubblicazione: (2025)
di: Atmer, Hannah, et al.
Pubblicazione: (2025)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
The Future of Memory: Limits and Opportunities
di: Dayo, Samuel, et al.
Pubblicazione: (2025)
di: Dayo, Samuel, et al.
Pubblicazione: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
di: Wang, Peipei, et al.
Pubblicazione: (2025)
di: Wang, Peipei, et al.
Pubblicazione: (2025)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
di: Arockiaraj, Jebacyril, et al.
Pubblicazione: (2025)
di: Arockiaraj, Jebacyril, et al.
Pubblicazione: (2025)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
di: Fu, Zizhuo, et al.
Pubblicazione: (2026)
di: Fu, Zizhuo, et al.
Pubblicazione: (2026)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
di: He, Zicheng, et al.
Pubblicazione: (2026)
di: He, Zicheng, et al.
Pubblicazione: (2026)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
di: Lee, Jungi, et al.
Pubblicazione: (2025)
di: Lee, Jungi, et al.
Pubblicazione: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
di: Zhang, Junming, et al.
Pubblicazione: (2026)
di: Zhang, Junming, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
di: Li, Jindong, et al.
Pubblicazione: (2025) -
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
di: Li, Jindong, et al.
Pubblicazione: (2024) -
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
di: Li, Tenglong, et al.
Pubblicazione: (2026) -
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
di: Li, Tenglong, et al.
Pubblicazione: (2024) -
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
di: Li, Tenglong, et al.
Pubblicazione: (2025)