ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Ziyu, Zhu, Tong, Zhang, Zhi, Fan, Tiantian, Yang, Jinluan, Kuang, Kun, Wei, Zhongyu, Wu, Fei, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
von: Shi, Ge, et al.
Veröffentlicht: (2025)
von: Shi, Ge, et al.
Veröffentlicht: (2025)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Discovering Invariant Neighborhood Patterns for Heterophilic Graphs
von: Yang, Jinluan, et al.
Veröffentlicht: (2024)
von: Yang, Jinluan, et al.
Veröffentlicht: (2024)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
von: Zhu, Xingkui, et al.
Veröffentlicht: (2024)
von: Zhu, Xingkui, et al.
Veröffentlicht: (2024)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
MoE-Loco: Mixture of Experts for Multitask Locomotion
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models
von: Tong, Yunze, et al.
Veröffentlicht: (2025)
von: Tong, Yunze, et al.
Veröffentlicht: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
Unifying Adversarial Perturbation for Graph Neural Networks
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2026)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
von: Yang, Jaewoo, et al.
Veröffentlicht: (2024)
DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
von: Feng, Yuchen, et al.
Veröffentlicht: (2025)
von: Feng, Yuchen, et al.
Veröffentlicht: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Do Domain-specific Experts exist in MoE-based LLMs?
von: Do, Giang, et al.
Veröffentlicht: (2026)
von: Do, Giang, et al.
Veröffentlicht: (2026)
MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices
von: Gavhane, Nishant, et al.
Veröffentlicht: (2025)
von: Gavhane, Nishant, et al.
Veröffentlicht: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
What Gets Activated: Uncovering Domain and Driver Experts in MoE Language Models
von: Hu, Guimin, et al.
Veröffentlicht: (2026)
von: Hu, Guimin, et al.
Veröffentlicht: (2026)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
von: Zheng, Youwei, et al.
Veröffentlicht: (2025)
von: Zheng, Youwei, et al.
Veröffentlicht: (2025)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025) -
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025) -
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
von: Shi, Ge, et al.
Veröffentlicht: (2025) -
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)