A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Wei-Hsing, Sharda, Janak, Shih, Cheng-Jhih, Kong, Yuyao, Waqar, Faaiq, Chen, Pin-Jun, Yingyan, Lin, Yu, Shimeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
von: Pan, Yue, et al.
Veröffentlicht: (2025)
von: Pan, Yue, et al.
Veröffentlicht: (2025)
3D-Carbon: An Analytical Carbon Modeling Tool for 3D and 2.5D Integrated Circuits
von: Zhao, Yujie, et al.
Veröffentlicht: (2023)
von: Zhao, Yujie, et al.
Veröffentlicht: (2023)
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2025)
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2025)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
Proxima: Near-storage Acceleration for Graph-based Approximate Nearest Neighbor Search in 3D NAND
von: Xu, Weihong, et al.
Veröffentlicht: (2023)
von: Xu, Weihong, et al.
Veröffentlicht: (2023)
CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms
von: Waqar, Faaiq, et al.
Veröffentlicht: (2025)
von: Waqar, Faaiq, et al.
Veröffentlicht: (2025)
Scalable Digital Compute-in-Memory Ising Machines for Robustness Verification of Binary Neural Networks
von: Vadlamani, Madhav, et al.
Veröffentlicht: (2026)
von: Vadlamani, Madhav, et al.
Veröffentlicht: (2026)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
von: Lee, Kiseok, et al.
Veröffentlicht: (2026)
von: Lee, Kiseok, et al.
Veröffentlicht: (2026)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
GauRast: Enhancing GPU Triangle Rasterizers to Accelerate 3D Gaussian Splatting
von: Li, Sixu, et al.
Veröffentlicht: (2025)
von: Li, Sixu, et al.
Veröffentlicht: (2025)
Instant-3D: Instant Neural Radiance Field Training Towards On-Device AR/VR 3D Reconstruction
von: Li, Sixu, et al.
Veröffentlicht: (2023)
von: Li, Sixu, et al.
Veröffentlicht: (2023)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
von: Hwang, Ranggi, et al.
Veröffentlicht: (2023)
von: Hwang, Ranggi, et al.
Veröffentlicht: (2023)
GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design
von: You, Haoran, et al.
Veröffentlicht: (2021)
von: You, Haoran, et al.
Veröffentlicht: (2021)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
von: Yi, Xiaoling, et al.
Veröffentlicht: (2026)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2026)
Analytical Heterogeneous Die-to-Die 3D Placement with Macros
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2024)
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
A Review of Multiscale Thermal Modeling in Heterogeneous 3D ICs
von: Barua, Baibhari Priya, et al.
Veröffentlicht: (2026)
von: Barua, Baibhari Priya, et al.
Veröffentlicht: (2026)
Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers
von: Xu, Boxun, et al.
Veröffentlicht: (2024)
von: Xu, Boxun, et al.
Veröffentlicht: (2024)
VersaQ-3D: A Reconfigurable Accelerator Enabling Feed-Forward and Generalizable 3D Reconstruction via Versatile Quantization
von: Zhang, Yipu, et al.
Veröffentlicht: (2026)
von: Zhang, Yipu, et al.
Veröffentlicht: (2026)
Spiking Transformer Hardware Accelerators in 3D Integration
von: Xu, Boxun, et al.
Veröffentlicht: (2024)
von: Xu, Boxun, et al.
Veröffentlicht: (2024)
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
ChatNeuroSim: An LLM Agent Framework for Automated Compute-in-Memory Accelerator Deployment and Optimization
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2026)
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2026)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
PC2IM: An Efficient In-Memory Computing Accelerator for 3D Point Cloud
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
von: Kim, Taehyun, et al.
Veröffentlicht: (2024)
von: Kim, Taehyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025) -
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024) -
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025) -
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
von: Pan, Yue, et al.
Veröffentlicht: (2025) -
3D-Carbon: An Analytical Carbon Modeling Tool for 3D and 2.5D Integrated Circuits
von: Zhao, Yujie, et al.
Veröffentlicht: (2023)