MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Taehyun, Choi, Kwanseok, Cho, Youngmock, Cho, Jaehoon, Lee, Hyuk-Jae, Sim, Jaewoong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
by: Lee, Jungi, et al.
Published: (2024)
by: Lee, Jungi, et al.
Published: (2024)
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
by: Yun, Sungmin, et al.
Published: (2024)
by: Yun, Sungmin, et al.
Published: (2024)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
by: Huang, Wei-Hsing, et al.
Published: (2025)
by: Huang, Wei-Hsing, et al.
Published: (2025)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
by: Yun, Sungmin, et al.
Published: (2025)
by: Yun, Sungmin, et al.
Published: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
by: Kim, Jungwoo, et al.
Published: (2026)
by: Kim, Jungwoo, et al.
Published: (2026)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
by: Cho, Hyunjun, et al.
Published: (2025)
by: Cho, Hyunjun, et al.
Published: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
by: Yoon, Jieon, et al.
Published: (2026)
by: Yoon, Jieon, et al.
Published: (2026)
VR-Pipe: Streamlining Hardware Graphics Pipeline for Volume Rendering
by: Lee, Junseo, et al.
Published: (2025)
by: Lee, Junseo, et al.
Published: (2025)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025)
by: Shin, Changmin, et al.
Published: (2025)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
by: Yoon, Hyunsung, et al.
Published: (2026)
by: Yoon, Hyunsung, et al.
Published: (2026)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
by: Cho, Eun-Su, et al.
Published: (2025)
by: Cho, Eun-Su, et al.
Published: (2025)
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
by: Cho, Sungjun, et al.
Published: (2024)
by: Cho, Sungjun, et al.
Published: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)
by: Heo, Guseul, et al.
Published: (2024)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
by: Kim, Youngsuk, et al.
Published: (2024)
by: Kim, Youngsuk, et al.
Published: (2024)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
by: Shende, Omkar B, et al.
Published: (2026)
by: Shende, Omkar B, et al.
Published: (2026)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
by: Han, Wontak, et al.
Published: (2024)
by: Han, Wontak, et al.
Published: (2024)
Subitizing-Inspired_Large_Language_Models_for_Floorplanning
by: Lu, Shao-Chien, et al.
Published: (2025)
by: Lu, Shao-Chien, et al.
Published: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
by: Gao, Hanyuan, et al.
Published: (2026)
by: Gao, Hanyuan, et al.
Published: (2026)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
by: Moon, Seungjae, et al.
Published: (2024)
by: Moon, Seungjae, et al.
Published: (2024)
GRTX: Efficient Ray Tracing for 3D Gaussian-Based Rendering
by: Lee, Junseo, et al.
Published: (2026)
by: Lee, Junseo, et al.
Published: (2026)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
by: Kyung, Kwanhee, et al.
Published: (2025)
by: Kyung, Kwanhee, et al.
Published: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
by: Lee, Kiseok, et al.
Published: (2026)
by: Lee, Kiseok, et al.
Published: (2026)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
by: Hwang, Soojin, et al.
Published: (2025)
by: Hwang, Soojin, et al.
Published: (2025)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
by: Lee, Minjae, et al.
Published: (2023)
by: Lee, Minjae, et al.
Published: (2023)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
Securing DRAM at Scale: ARFM-Driven Row Hammer Defense with Unveiling the Threat of Short tRC Patterns
by: Joo, Nogeun, et al.
Published: (2025)
by: Joo, Nogeun, et al.
Published: (2025)
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
by: Xu, Weikai, et al.
Published: (2026)
by: Xu, Weikai, et al.
Published: (2026)
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
by: Jang, Hongsun, et al.
Published: (2024)
by: Jang, Hongsun, et al.
Published: (2024)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
by: Shin, Hery, et al.
Published: (2024)
by: Shin, Hery, et al.
Published: (2024)
LASANA: Large-Scale Surrogate Modeling for Analog Neuromorphic Architecture Exploration
by: Ho, Jason, et al.
Published: (2025)
by: Ho, Jason, et al.
Published: (2025)
Similar Items
-
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025) -
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
by: Lee, Jungi, et al.
Published: (2024) -
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
by: Yun, Sungmin, et al.
Published: (2024) -
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
by: Huang, Wei-Hsing, et al.
Published: (2025) -
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)