A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Cong, Xue, Chenhao, Ren, Yi, Dong, Xiping, Cheng, Yu, Hu, Yinbo, Bai, Fujun, Guo, Yixin, Jiang, Xiping, Wu, Qiang, Yang, Zhi, Cheng, Zhe, Xie, Yuan, Sun, Guangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
von: Chen, Yiqi, et al.
Veröffentlicht: (2025)
von: Chen, Yiqi, et al.
Veröffentlicht: (2025)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
von: Canpolat, Oğuzhan, et al.
Veröffentlicht: (2025)
von: Canpolat, Oğuzhan, et al.
Veröffentlicht: (2025)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
von: Olgun, Ataberk, et al.
Veröffentlicht: (2022)
von: Olgun, Ataberk, et al.
Veröffentlicht: (2022)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips
von: Olgun, Ataberk, et al.
Veröffentlicht: (2022)
von: Olgun, Ataberk, et al.
Veröffentlicht: (2022)
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
Shifting in-DRAM
von: Tegge, William C., et al.
Veröffentlicht: (2026)
von: Tegge, William C., et al.
Veröffentlicht: (2026)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
von: Cai, Jingwei, et al.
Veröffentlicht: (2025)
von: Cai, Jingwei, et al.
Veröffentlicht: (2025)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
von: Cai, Victor, et al.
Veröffentlicht: (2025)
von: Cai, Victor, et al.
Veröffentlicht: (2025)
EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
von: Huang, Yi, et al.
Veröffentlicht: (2025)
von: Huang, Yi, et al.
Veröffentlicht: (2025)
Exploring DRAM Cache Prefetching for Pooled Memory
von: Tirumalasetty, Chandrahas, et al.
Veröffentlicht: (2024)
von: Tirumalasetty, Chandrahas, et al.
Veröffentlicht: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
Rethinking the Producer-Consumer Relationship in Modern DRAM-Based Systems
von: Patel, Minesh, et al.
Veröffentlicht: (2024)
von: Patel, Minesh, et al.
Veröffentlicht: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/Write
von: Guo, Yan-Cheng, et al.
Veröffentlicht: (2026)
von: Guo, Yan-Cheng, et al.
Veröffentlicht: (2026)
ARTEMIS: A Mixed Analog-Stochastic In-DRAM Accelerator for Transformer Neural Networks
von: Afifi, Salma, et al.
Veröffentlicht: (2024)
von: Afifi, Salma, et al.
Veröffentlicht: (2024)
DRAM-Profiler: An Experimental DRAM RowHammer Vulnerability Profiling Mechanism
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
Beehive: A Flexible Network Stack for Direct-Attached Accelerators
von: Lim, Katie, et al.
Veröffentlicht: (2024)
von: Lim, Katie, et al.
Veröffentlicht: (2024)
3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event Cameras
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
Corrigendum to: A Systematic Study of DDR4 DRAM Faults in the Field
von: Beigi, Majed Valad, et al.
Veröffentlicht: (2024)
von: Beigi, Majed Valad, et al.
Veröffentlicht: (2024)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators
von: Yik, Jason, et al.
Veröffentlicht: (2025)
von: Yik, Jason, et al.
Veröffentlicht: (2025)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
von: Cho, Hyunjun, et al.
Veröffentlicht: (2025)
von: Cho, Hyunjun, et al.
Veröffentlicht: (2025)
Self-Managing DRAM: A Low-Cost Framework for Enabling Autonomous and Efficient in-DRAM Operations
von: Hassan, Hasan, et al.
Veröffentlicht: (2022)
von: Hassan, Hasan, et al.
Veröffentlicht: (2022)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
von: Agosta, Giovanni, et al.
Veröffentlicht: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026) -
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
von: Chen, Yiqi, et al.
Veröffentlicht: (2025) -
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026) -
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025) -
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
von: Canpolat, Oğuzhan, et al.
Veröffentlicht: (2025)