FloE: On-the-Fly MoE Inference on Memory-constrained GPU
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yuxin, Li, Zheng, Zhang, Jun, Wang, Jue, Wang, Yiping, Xie, Zhongle, Chen, Ke, Shou, Lidan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Shapley Value in Data Analytics
von: Lin, Hong, et al.
Veröffentlicht: (2024)
von: Lin, Hong, et al.
Veröffentlicht: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution
von: Son, Muyoung, et al.
Veröffentlicht: (2026)
von: Son, Muyoung, et al.
Veröffentlicht: (2026)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data Warehouses
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
von: Peng, Yuchen, et al.
Veröffentlicht: (2026)
von: Peng, Yuchen, et al.
Veröffentlicht: (2026)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
von: Huang, Yuegui, et al.
Veröffentlicht: (2026)
von: Huang, Yuegui, et al.
Veröffentlicht: (2026)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
von: Zhu, Xiongwei, et al.
Veröffentlicht: (2026)
von: Zhu, Xiongwei, et al.
Veröffentlicht: (2026)
Faster MoE LLM Inference for Extremely Large Models
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
Making MoE-based LLM Inference Resilient with Tarragon
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models
von: Jin, Lyudong, et al.
Veröffentlicht: (2025)
von: Jin, Lyudong, et al.
Veröffentlicht: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
GRIN: GRadient-INformed MoE
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Fast MoE Inference via Predictive Prefetching and Expert Replication
von: Jyothish, Ankit, et al.
Veröffentlicht: (2026)
von: Jyothish, Ankit, et al.
Veröffentlicht: (2026)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
von: Hu, Tianlun, et al.
Veröffentlicht: (2026)
von: Hu, Tianlun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning
von: Zhang, Jun, et al.
Veröffentlicht: (2025) -
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024) -
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025) -
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025) -
A Comprehensive Study of Shapley Value in Data Analytics
von: Lin, Hong, et al.
Veröffentlicht: (2024)