Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jungwoo, Lacouture, Rubens, Zhang, Genghan, Sohn, Gina, Zhang, Qizheng, Gandhi, Swapnil, Kozyrakis, Christos, Olukotun, Kunle |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024)
por: Sohn, Gina, et al.
Publicado: (2024)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
por: Sohn, Gina, et al.
Publicado: (2025)
por: Sohn, Gina, et al.
Publicado: (2025)
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
por: Ko, Sho, et al.
Publicado: (2025)
por: Ko, Sho, et al.
Publicado: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
por: Zhou, Zikai, et al.
Publicado: (2025)
por: Zhou, Zikai, et al.
Publicado: (2025)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025)
por: Lacouture, Rubens, et al.
Publicado: (2025)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
por: Ko, Sho, et al.
Publicado: (2024)
por: Ko, Sho, et al.
Publicado: (2024)
AME-PIM: Can Memory be Your Next Tensor Accelerator?
por: Venieri, Emanuele, et al.
Publicado: (2026)
por: Venieri, Emanuele, et al.
Publicado: (2026)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
por: Lee, Dongjae, et al.
Publicado: (2024)
por: Lee, Dongjae, et al.
Publicado: (2024)
Revet: A Language and Compiler for Dataflow Threads
por: Rucker, Alexander, et al.
Publicado: (2023)
por: Rucker, Alexander, et al.
Publicado: (2023)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
por: Alsop, Johnathan, et al.
Publicado: (2023)
por: Alsop, Johnathan, et al.
Publicado: (2023)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
por: Davies, Michael, et al.
Publicado: (2025)
por: Davies, Michael, et al.
Publicado: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
por: Seo, Minseok, et al.
Publicado: (2024)
por: Seo, Minseok, et al.
Publicado: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
por: Hyun, Bongjoon, et al.
Publicado: (2023)
por: Hyun, Bongjoon, et al.
Publicado: (2023)
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
por: Yun, Sungmin, et al.
Publicado: (2024)
por: Yun, Sungmin, et al.
Publicado: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
por: Heo, Guseul, et al.
Publicado: (2024)
por: Heo, Guseul, et al.
Publicado: (2024)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
por: Dong, Jiale, et al.
Publicado: (2025)
por: Dong, Jiale, et al.
Publicado: (2025)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
por: Kim, Youngsuk, et al.
Publicado: (2024)
por: Kim, Youngsuk, et al.
Publicado: (2024)
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
por: Cha, SangHoon, et al.
Publicado: (2026)
por: Cha, SangHoon, et al.
Publicado: (2026)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
por: Lee, Dongjae, et al.
Publicado: (2025)
por: Lee, Dongjae, et al.
Publicado: (2025)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
por: Huang, Wei-Hsing, et al.
Publicado: (2025)
por: Huang, Wei-Hsing, et al.
Publicado: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
por: Gao, Hanyuan, et al.
Publicado: (2026)
por: Gao, Hanyuan, et al.
Publicado: (2026)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
por: Lin, Ye, et al.
Publicado: (2026)
por: Lin, Ye, et al.
Publicado: (2026)
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
por: Shekar, Akhil, et al.
Publicado: (2025)
por: Shekar, Akhil, et al.
Publicado: (2025)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
por: Wu, Yuting, et al.
Publicado: (2023)
por: Wu, Yuting, et al.
Publicado: (2023)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
por: Kyung, Kwanhee, et al.
Publicado: (2025)
por: Kyung, Kwanhee, et al.
Publicado: (2025)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
por: Ma, Songchen, et al.
Publicado: (2026)
por: Ma, Songchen, et al.
Publicado: (2026)
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
por: Yun, Sungmin, et al.
Publicado: (2025)
por: Yun, Sungmin, et al.
Publicado: (2025)
Annotated PIM Bibliography
por: Kogge, Peter M.
Publicado: (2026)
por: Kogge, Peter M.
Publicado: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
por: Dong, Jiale, et al.
Publicado: (2025)
por: Dong, Jiale, et al.
Publicado: (2025)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
por: Tajdari, Sabiha, et al.
Publicado: (2025)
por: Tajdari, Sabiha, et al.
Publicado: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
por: Hong, Junguk, et al.
Publicado: (2026)
por: Hong, Junguk, et al.
Publicado: (2026)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
por: Wang, Xuan, et al.
Publicado: (2024)
por: Wang, Xuan, et al.
Publicado: (2024)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
por: Jeon, Sangmin, et al.
Publicado: (2025)
por: Jeon, Sangmin, et al.
Publicado: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
por: Kanani, Alish, et al.
Publicado: (2025)
por: Kanani, Alish, et al.
Publicado: (2025)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
por: Kim, Jae-Young, et al.
Publicado: (2025)
por: Kim, Jae-Young, et al.
Publicado: (2025)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
por: Xu, Weikai, et al.
Publicado: (2026)
por: Xu, Weikai, et al.
Publicado: (2026)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
por: Han, Wontak, et al.
Publicado: (2024)
por: Han, Wontak, et al.
Publicado: (2024)
Ejemplares similares
-
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024) -
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
por: Sohn, Gina, et al.
Publicado: (2025) -
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
por: Ko, Sho, et al.
Publicado: (2025) -
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
por: Zhou, Zikai, et al.
Publicado: (2025) -
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025)