Salvato in:
| Autori principali: | Ma, Haiyue, Du, Zhixu, Chen, Yiran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.07366 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlashMoE: Fast Distributed MoE in a Single Kernel
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
di: Choi, Yuseon, et al.
Pubblicazione: (2025)
di: Choi, Yuseon, et al.
Pubblicazione: (2025)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
di: Pan, Yue, et al.
Pubblicazione: (2025)
di: Pan, Yue, et al.
Pubblicazione: (2025)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
di: Ma, Songchen, et al.
Pubblicazione: (2026)
di: Ma, Songchen, et al.
Pubblicazione: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
di: Shende, Omkar B, et al.
Pubblicazione: (2026)
di: Shende, Omkar B, et al.
Pubblicazione: (2026)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
di: Yu, Zhongkai, et al.
Pubblicazione: (2025)
di: Yu, Zhongkai, et al.
Pubblicazione: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
di: Luo, Shuqing, et al.
Pubblicazione: (2026)
di: Luo, Shuqing, et al.
Pubblicazione: (2026)
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
di: Choi, Yuseon, et al.
Pubblicazione: (2026)
di: Choi, Yuseon, et al.
Pubblicazione: (2026)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
di: Lv, Bo, et al.
Pubblicazione: (2026)
di: Lv, Bo, et al.
Pubblicazione: (2026)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
On the Shape of Latent Variables in a Denoising VAE-MoG: A Posterior Sampling-Based Study
di: Bascuñán, Fernanda Zapata
Pubblicazione: (2025)
di: Bascuñán, Fernanda Zapata
Pubblicazione: (2025)
Accelerating Frontier MoE Training with 3D Integrated Optics
di: Bernadskiy, Mikhail, et al.
Pubblicazione: (2025)
di: Bernadskiy, Mikhail, et al.
Pubblicazione: (2025)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
di: Yang, Xiaoxuan, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxuan, et al.
Pubblicazione: (2025)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
di: Dong, Jiale, et al.
Pubblicazione: (2025)
di: Dong, Jiale, et al.
Pubblicazione: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
di: Duan, Bowen, et al.
Pubblicazione: (2026)
di: Duan, Bowen, et al.
Pubblicazione: (2026)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
di: Fan, Zehao, et al.
Pubblicazione: (2025)
di: Fan, Zehao, et al.
Pubblicazione: (2025)
Graph Neural Networks Based Analog Circuit Link Prediction
di: Pan, Guanyuan, et al.
Pubblicazione: (2025)
di: Pan, Guanyuan, et al.
Pubblicazione: (2025)
Unsupervised Graph Neural Network Framework for Balanced Multipatterning in Advanced Electronic Design Automation Layouts
di: Helaly, Abdelrahman, et al.
Pubblicazione: (2025)
di: Helaly, Abdelrahman, et al.
Pubblicazione: (2025)
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
di: Yun, Sungmin, et al.
Pubblicazione: (2024)
di: Yun, Sungmin, et al.
Pubblicazione: (2024)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
di: Chowdhury, Md Rownak Hossain, et al.
Pubblicazione: (2025)
di: Chowdhury, Md Rownak Hossain, et al.
Pubblicazione: (2025)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
di: Molom-Ochir, Tergel, et al.
Pubblicazione: (2024)
di: Molom-Ochir, Tergel, et al.
Pubblicazione: (2024)
CAMformer: Associative Memory is All You Need
di: Molom-Ochir, Tergel, et al.
Pubblicazione: (2025)
di: Molom-Ochir, Tergel, et al.
Pubblicazione: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology Generation
di: Chang, Chen-Chia, et al.
Pubblicazione: (2025)
di: Chang, Chen-Chia, et al.
Pubblicazione: (2025)
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
di: Alama, Youssef A. Ait, et al.
Pubblicazione: (2024)
di: Alama, Youssef A. Ait, et al.
Pubblicazione: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
di: Sharma, Amit
Pubblicazione: (2025)
di: Sharma, Amit
Pubblicazione: (2025)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
di: Liu, Qingyuan, et al.
Pubblicazione: (2025)
di: Liu, Qingyuan, et al.
Pubblicazione: (2025)
PICBench: Benchmarking LLMs for Photonic Integrated Circuits Design
di: Wu, Yuchao, et al.
Pubblicazione: (2025)
di: Wu, Yuchao, et al.
Pubblicazione: (2025)
LaMAGIC: Language-Model-based Topology Generation for Analog Integrated Circuits
di: Chang, Chen-Chia, et al.
Pubblicazione: (2024)
di: Chang, Chen-Chia, et al.
Pubblicazione: (2024)
PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models
di: Gupta, Neelesh, et al.
Pubblicazione: (2024)
di: Gupta, Neelesh, et al.
Pubblicazione: (2024)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
di: Mao, Gang, et al.
Pubblicazione: (2025)
di: Mao, Gang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FlashMoE: Fast Distributed MoE in a Single Kernel
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025) -
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025) -
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
di: Hwang, Ranggi, et al.
Pubblicazione: (2023) -
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
di: Choi, Yuseon, et al.
Pubblicazione: (2025) -
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
di: Pan, Yue, et al.
Pubblicazione: (2025)