Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinyu, Li, Jieyu, Sun, Yanan, He, Weifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
PC2IM: An Efficient In-Memory Computing Accelerator for 3D Point Cloud
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
von: Zhang, Jie, et al.
Veröffentlicht: (2026)
von: Zhang, Jie, et al.
Veröffentlicht: (2026)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
In-Memory Computing Architecture for Efficient Hardware Security
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
von: He, Zifan, et al.
Veröffentlicht: (2025)
von: He, Zifan, et al.
Veröffentlicht: (2025)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
von: Wang, Kainan, et al.
Veröffentlicht: (2025)
von: Wang, Kainan, et al.
Veröffentlicht: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
von: You, Dean, et al.
Veröffentlicht: (2025)
von: You, Dean, et al.
Veröffentlicht: (2025)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
von: Wan, Zishen, et al.
Veröffentlicht: (2025)
von: Wan, Zishen, et al.
Veröffentlicht: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
von: Ai, Chenyang, et al.
Veröffentlicht: (2026)
von: Ai, Chenyang, et al.
Veröffentlicht: (2026)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
von: Sun, He, et al.
Veröffentlicht: (2026)
von: Sun, He, et al.
Veröffentlicht: (2026)
SCREME: A Scalable Framework for Resilient Memory Design
von: Li, Fan, et al.
Veröffentlicht: (2025)
von: Li, Fan, et al.
Veröffentlicht: (2025)
Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
von: Pan, Tianwei, et al.
Veröffentlicht: (2025)
von: Pan, Tianwei, et al.
Veröffentlicht: (2025)
When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge Computing
von: Ren, Haoxiong, et al.
Veröffentlicht: (2025)
von: Ren, Haoxiong, et al.
Veröffentlicht: (2025)
ARMOR: Robust and Efficient CNN-Based SAR ATR through Model-Hardware Co-Design
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2026)
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2026)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
von: He, Mingxuan, et al.
Veröffentlicht: (2024)
von: He, Mingxuan, et al.
Veröffentlicht: (2024)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
von: Tan, Zhanhong, et al.
Veröffentlicht: (2024)
von: Tan, Zhanhong, et al.
Veröffentlicht: (2024)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
von: Chen, Jinwu, et al.
Veröffentlicht: (2026)
von: Chen, Jinwu, et al.
Veröffentlicht: (2026)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
von: Manca, Federico, et al.
Veröffentlicht: (2024)
von: Manca, Federico, et al.
Veröffentlicht: (2024)
From Natural Language to Silicon: The Representation Bottleneck in LLM Hardware Design
von: Fu, Weimin, et al.
Veröffentlicht: (2026)
von: Fu, Weimin, et al.
Veröffentlicht: (2026)
Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection
von: Kim, Junhwan, et al.
Veröffentlicht: (2026)
von: Kim, Junhwan, et al.
Veröffentlicht: (2026)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025) -
Pushing the Limits of BFP on Narrow Precision LLM Inference
von: Wang, Hui, et al.
Veröffentlicht: (2025) -
PC2IM: An Efficient In-Memory Computing Accelerator for 3D Point Cloud
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026) -
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
von: Zhang, Jie, et al.
Veröffentlicht: (2026) -
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)