Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Cong, Yin, Yihan, Xue, Chenhao, Wang, Zhao, Bai, Fujun, Guo, Yixin, Jiang, Xiping, Wu, Qiang, Xie, Yuan, Sun, Guangyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
di: Li, Cong, et al.
Pubblicazione: (2026)
di: Li, Cong, et al.
Pubblicazione: (2026)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2026)
di: Lu, Tsung-Han, et al.
Pubblicazione: (2026)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021)
di: Wang, Zhao, et al.
Pubblicazione: (2021)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
di: Xue, Chenhao, et al.
Pubblicazione: (2026)
di: Xue, Chenhao, et al.
Pubblicazione: (2026)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2025)
di: Lu, Tsung-Han, et al.
Pubblicazione: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
di: Alsop, Johnathan, et al.
Pubblicazione: (2023)
di: Alsop, Johnathan, et al.
Pubblicazione: (2023)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
di: Zhang, Yuanpeng, et al.
Pubblicazione: (2025)
di: Zhang, Yuanpeng, et al.
Pubblicazione: (2025)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
di: Kim, Dong Eun, et al.
Pubblicazione: (2025)
di: Kim, Dong Eun, et al.
Pubblicazione: (2025)
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025)
di: Ren, Yi, et al.
Pubblicazione: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
di: Liao, Yuchao, et al.
Pubblicazione: (2026)
di: Liao, Yuchao, et al.
Pubblicazione: (2026)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
di: Tsai, Pei-Huan, et al.
Pubblicazione: (2026)
di: Tsai, Pei-Huan, et al.
Pubblicazione: (2026)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
di: Pan, Yue, et al.
Pubblicazione: (2025)
di: Pan, Yue, et al.
Pubblicazione: (2025)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025)
di: Ren, Yi, et al.
Pubblicazione: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
di: You, Dean, et al.
Pubblicazione: (2025)
di: You, Dean, et al.
Pubblicazione: (2025)
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
di: Shekar, Akhil, et al.
Pubblicazione: (2025)
di: Shekar, Akhil, et al.
Pubblicazione: (2025)
Shifting in-DRAM
di: Tegge, William C., et al.
Pubblicazione: (2026)
di: Tegge, William C., et al.
Pubblicazione: (2026)
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
di: Xuan, Wei, et al.
Pubblicazione: (2025)
di: Xuan, Wei, et al.
Pubblicazione: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
di: Wang, Peipei, et al.
Pubblicazione: (2025)
di: Wang, Peipei, et al.
Pubblicazione: (2025)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
di: Zou, Yuyang, et al.
Pubblicazione: (2025)
di: Zou, Yuyang, et al.
Pubblicazione: (2025)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
di: Kiyawat, Khyati, et al.
Pubblicazione: (2025)
di: Kiyawat, Khyati, et al.
Pubblicazione: (2025)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
di: Wang, Wenxun, et al.
Pubblicazione: (2025)
di: Wang, Wenxun, et al.
Pubblicazione: (2025)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
di: Cai, Jingwei, et al.
Pubblicazione: (2025)
di: Cai, Jingwei, et al.
Pubblicazione: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
NDFT: Accelerating Density Functional Theory Calculations via Hardware/Software Co-Design on Near-Data Computing System
di: Jiang, Qingcai, et al.
Pubblicazione: (2025)
di: Jiang, Qingcai, et al.
Pubblicazione: (2025)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
di: Olgun, Ataberk, et al.
Pubblicazione: (2022)
di: Olgun, Ataberk, et al.
Pubblicazione: (2022)
FireBridge: Cycle-Accurate Hardware + Firmware Co-Verification for Modern Accelerators
di: Abarajithan, G, et al.
Pubblicazione: (2026)
di: Abarajithan, G, et al.
Pubblicazione: (2026)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
di: Chen, Jinwu, et al.
Pubblicazione: (2026)
di: Chen, Jinwu, et al.
Pubblicazione: (2026)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
di: Li, Boyu, et al.
Pubblicazione: (2025)
di: Li, Boyu, et al.
Pubblicazione: (2025)
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
di: Canpolat, Oğuzhan, et al.
Pubblicazione: (2025)
di: Canpolat, Oğuzhan, et al.
Pubblicazione: (2025)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
di: Zou, Guoqiang, et al.
Pubblicazione: (2025)
di: Zou, Guoqiang, et al.
Pubblicazione: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAs
di: Lee, Jiwoon, et al.
Pubblicazione: (2026)
di: Lee, Jiwoon, et al.
Pubblicazione: (2026)
Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
di: Pan, Tianwei, et al.
Pubblicazione: (2025)
di: Pan, Tianwei, et al.
Pubblicazione: (2025)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
di: Zhang, Chenguang, et al.
Pubblicazione: (2024)
di: Zhang, Chenguang, et al.
Pubblicazione: (2024)
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
di: Bazzi, Jinane, et al.
Pubblicazione: (2026)
di: Bazzi, Jinane, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
di: Li, Cong, et al.
Pubblicazione: (2026) -
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2026) -
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021) -
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
di: Xue, Chenhao, et al.
Pubblicazione: (2026) -
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2025)