Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Yue, Xia, Zihan, Hsu, Po-Kai, Hu, Lanxiang, Kim, Hyungyo, Sharda, Janak, Zhou, Minxuan, Kim, Nam Sung, Yu, Shimeng, Rosing, Tajana, Kang, Mingu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914078827479040
author Pan, Yue
Xia, Zihan
Hsu, Po-Kai
Hu, Lanxiang
Kim, Hyungyo
Sharda, Janak
Zhou, Minxuan
Kim, Nam Sung
Yu, Shimeng
Rosing, Tajana
Kang, Mingu
author_facet Pan, Yue
Xia, Zihan
Hsu, Po-Kai
Hu, Lanxiang
Kim, Hyungyo
Sharda, Janak
Zhou, Minxuan
Kim, Nam Sung
Yu, Shimeng
Rosing, Tajana
Kang, Mingu
contents As Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks. MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models. However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers. To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration. The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer. Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing. Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the z-dimension by constructing internal memory tiers and assigning data across layers based on access likelihood, guided by topic-based expert usage prediction to boost NMP throughput. The Stratum system achieves up to 8.29x improvement in decoding throughput and 7.66x better energy efficiency across various benchmarks compared to GPU baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
Pan, Yue
Xia, Zihan
Hsu, Po-Kai
Hu, Lanxiang
Kim, Hyungyo
Sharda, Janak
Zhou, Minxuan
Kim, Nam Sung
Yu, Shimeng
Rosing, Tajana
Kang, Mingu
Hardware Architecture
Emerging Technologies
Machine Learning
As Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks. MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models. However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers. To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration. The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer. Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing. Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the z-dimension by constructing internal memory tiers and assigning data across layers based on access likelihood, guided by topic-based expert usage prediction to boost NMP throughput. The Stratum system achieves up to 8.29x improvement in decoding throughput and 7.66x better energy efficiency across various benchmarks compared to GPU baselines.
title Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
topic Hardware Architecture
Emerging Technologies
Machine Learning
url https://arxiv.org/abs/2510.05245