Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yinghan, Li, Yifei, Zhang, Jiejing, Chen, Bujiao, Chen, Xiaotong, Duan, Lian, Jin, Yejun, Li, Zheng, Liu, Xuanyu, Wang, Haoyu, Wang, Wente, Wang, Yajie, Yang, Jiacheng, Zhang, Peiyang, Zheng, Laiwen, Yu, Wenyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RadiK: Scalable and Optimized GPU-Parallel Radix Top-K Selection
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
What Are "Adult Services"?
von: Wente, Joyce
Veröffentlicht: (1979)
von: Wente, Joyce
Veröffentlicht: (1979)
FloE: On-the-Fly MoE Inference on Memory-constrained GPU
von: Zhou, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhou, Yuxin, et al.
Veröffentlicht: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
TRUST: Triangle Counting Reloaded on GPUs
von: Pandey, Santosh, et al.
Veröffentlicht: (2021)
von: Pandey, Santosh, et al.
Veröffentlicht: (2021)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
Magnetic helicity evolution during active region emergence and subsequent flare productivity
von: Sun, Zheng, et al.
Veröffentlicht: (2024)
von: Sun, Zheng, et al.
Veröffentlicht: (2024)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
von: Xu, Yechen, et al.
Veröffentlicht: (2026)
von: Xu, Yechen, et al.
Veröffentlicht: (2026)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Effects of the Grazing Exclusion Policy on Pheasant Nesting Success and Predation Risk in the Protected Areas of Southwestern China
von: Yifei Zhang, et al.
Veröffentlicht: (2025)
von: Yifei Zhang, et al.
Veröffentlicht: (2025)
Batch Array Codes
von: Kong, Xiangliang, et al.
Veröffentlicht: (2024)
von: Kong, Xiangliang, et al.
Veröffentlicht: (2024)
LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers
von: Yu, Enda, et al.
Veröffentlicht: (2025)
von: Yu, Enda, et al.
Veröffentlicht: (2025)
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
rLLM: Relational Table Learning with LLMs
von: Li, Weichen, et al.
Veröffentlicht: (2024)
von: Li, Weichen, et al.
Veröffentlicht: (2024)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
Comparative Efficacy of Various Interventions to Reduce Perceived Stress Among Older Adults: A Systematic Review and Network Meta‐Analysis
von: Mingyue Zhu, et al.
Veröffentlicht: (2025)
von: Mingyue Zhu, et al.
Veröffentlicht: (2025)
NESA: Relational Neuro-Symbolic Static Program Analysis
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
von: Wang, Chengpeng, et al.
Veröffentlicht: (2024)
Irregular threefolds with numerically trivial canonical divisor
von: Chen, Jingshan, et al.
Veröffentlicht: (2024)
von: Chen, Jingshan, et al.
Veröffentlicht: (2024)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
Can BEV Perception Gracefully Degrade under Sensor Failures?
von: Zhang, Haifa, et al.
Veröffentlicht: (2026)
von: Zhang, Haifa, et al.
Veröffentlicht: (2026)
Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Strategic Joining and Optimal Pricing in a Single‐Server Batch Arrival Queue With Different Information of Batch Size
von: Kaili Li, et al.
Veröffentlicht: (2025)
von: Kaili Li, et al.
Veröffentlicht: (2025)
Leakage Benchmarking for Universal Gate Sets
von: Wu, Bujiao, et al.
Veröffentlicht: (2023)
von: Wu, Bujiao, et al.
Veröffentlicht: (2023)
Adaptive-depth randomized measurement for fermionic observables
von: Bian, Kaiming, et al.
Veröffentlicht: (2025)
von: Bian, Kaiming, et al.
Veröffentlicht: (2025)
A Practical Introduction to Deep Reinforcement Learning
von: Sun, Yinghan, et al.
Veröffentlicht: (2025)
von: Sun, Yinghan, et al.
Veröffentlicht: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
GenTS: A Comprehensive Benchmark Library for Generative Time Series Models
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RadiK: Scalable and Optimized GPU-Parallel Radix Top-K Selection
von: Li, Yifei, et al.
Veröffentlicht: (2025) -
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026) -
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025) -
What Are "Adult Services"?
von: Wente, Joyce
Veröffentlicht: (1979) -
FloE: On-the-Fly MoE Inference on Memory-constrained GPU
von: Zhou, Yuxin, et al.
Veröffentlicht: (2025)