PreMoE: Proactive Inference for Efficient Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Pei, Zehua, Zhang, Ying, Zhen, Hui-Ling, Yuan, Tao, Yu, Xianzhi, Dong, Zhenhua, Pan, Sinno Jialin, Yuan, Mingxuan, Yu, Bei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
by: Pei, Zehua, et al.
Published: (2026)
by: Pei, Zehua, et al.
Published: (2026)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation
by: Pei, Zehua, et al.
Published: (2026)
by: Pei, Zehua, et al.
Published: (2026)
SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
MemDLM: Memory-Enhanced DLM Training
by: Pei, Zehua, et al.
Published: (2026)
by: Pei, Zehua, et al.
Published: (2026)
Behavioral Fingerprinting of Large Language Models
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
BetterV: Controlled Verilog Generation with Discriminative Guidance
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
by: Jiang, Weisen, et al.
Published: (2026)
by: Jiang, Weisen, et al.
Published: (2026)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
by: Liu, Zehua, et al.
Published: (2025)
by: Liu, Zehua, et al.
Published: (2025)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
by: Xing, Zeyu, et al.
Published: (2026)
by: Xing, Zeyu, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
DiLA: Enhancing LLM Tool Learning with Differential Logic Layer
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
by: Tian, Anxin, et al.
Published: (2026)
by: Tian, Anxin, et al.
Published: (2026)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
by: Yankun, Hong, et al.
Published: (2025)
by: Yankun, Hong, et al.
Published: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
by: Wang, Wenfeng, et al.
Published: (2025)
by: Wang, Wenfeng, et al.
Published: (2025)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
by: Lin, Haoran, et al.
Published: (2025)
by: Lin, Haoran, et al.
Published: (2025)
Learning Gradient-based Mixup with Extrapolation toward Flatter Minima for Domain Generalization
by: Peng, Danni, et al.
Published: (2022)
by: Peng, Danni, et al.
Published: (2022)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
by: Jiang, Weisen, et al.
Published: (2025)
by: Jiang, Weisen, et al.
Published: (2025)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
Towards Efficient Agents: A Co-Design of Inference Architecture and System
by: Lin, Weizhe, et al.
Published: (2025)
by: Lin, Weizhe, et al.
Published: (2025)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
by: Sun, Shengyin, et al.
Published: (2026)
by: Sun, Shengyin, et al.
Published: (2026)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
by: Singh, Gursimran, et al.
Published: (2025)
by: Singh, Gursimran, et al.
Published: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
by: Lin, Weizhe, et al.
Published: (2025)
by: Lin, Weizhe, et al.
Published: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
by: Liu, Zehua, et al.
Published: (2026)
by: Liu, Zehua, et al.
Published: (2026)
Fast Graph Generation via Spectral Diffusion
by: Luo, Tianze, et al.
Published: (2022)
by: Luo, Tianze, et al.
Published: (2022)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
by: Wu, Han, et al.
Published: (2025)
by: Wu, Han, et al.
Published: (2025)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
Similar Items
-
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
by: Pei, Zehua, et al.
Published: (2026) -
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
by: Pei, Zehua, et al.
Published: (2024) -
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
by: Pei, Zehua, et al.
Published: (2025) -
Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation
by: Pei, Zehua, et al.
Published: (2026) -
SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
by: Pei, Zehua, et al.
Published: (2025)