Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Hanyuan, Yang, Xiaoxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
von: Dong, Shuai, et al.
Veröffentlicht: (2026)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
Exploring DRAM Cache Prefetching for Pooled Memory
von: Tirumalasetty, Chandrahas, et al.
Veröffentlicht: (2024)
von: Tirumalasetty, Chandrahas, et al.
Veröffentlicht: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
Improving the Representativeness of Simulation Intervals for the Cache Memory System
von: Bueno, Nicolas, et al.
Veröffentlicht: (2024)
von: Bueno, Nicolas, et al.
Veröffentlicht: (2024)
In-Memory Computing Architecture for Efficient Hardware Security
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
von: Dong, Jiale, et al.
Veröffentlicht: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
Functionality Locality, Mixture & Control = Logic = Memory
von: Peng, Xiangjun
Veröffentlicht: (2024)
von: Peng, Xiangjun
Veröffentlicht: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
In-place Switch: Reprogramming based SLC Cache Design for Hybrid 3D SSDs
von: Yang, Xufeng, et al.
Veröffentlicht: (2024)
von: Yang, Xufeng, et al.
Veröffentlicht: (2024)
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
von: Khadem, Alireza, et al.
Veröffentlicht: (2025)
von: Khadem, Alireza, et al.
Veröffentlicht: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
von: You, Dean, et al.
Veröffentlicht: (2025)
von: You, Dean, et al.
Veröffentlicht: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
Rhea: a Framework for Fast Design and Validation of RTL Cache-Coherent Memory Subsystems
von: Zoni, Davide, et al.
Veröffentlicht: (2025)
von: Zoni, Davide, et al.
Veröffentlicht: (2025)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
von: Vafaei, Jafar, et al.
Veröffentlicht: (2024)
von: Vafaei, Jafar, et al.
Veröffentlicht: (2024)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
PC2IM: An Efficient In-Memory Computing Accelerator for 3D Point Cloud
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
von: Wang, Dengfeng, et al.
Veröffentlicht: (2026)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache
von: Cheshmikhani, Elham, et al.
Veröffentlicht: (2025)
von: Cheshmikhani, Elham, et al.
Veröffentlicht: (2025)
An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision Sensors
von: Zhao, Qinghang, et al.
Veröffentlicht: (2024)
von: Zhao, Qinghang, et al.
Veröffentlicht: (2024)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms
von: Khan, Asif Ali, et al.
Veröffentlicht: (2022)
von: Khan, Asif Ali, et al.
Veröffentlicht: (2022)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
Voxel-CIM: An Efficient Compute-in-Memory Accelerator for Voxel-based Point Cloud Neural Networks
von: Lin, Xipeng, et al.
Veröffentlicht: (2024)
von: Lin, Xipeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
von: Chen, Peilin, et al.
Veröffentlicht: (2025) -
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
von: Chen, Peilin, et al.
Veröffentlicht: (2025) -
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024) -
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024) -
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2026)