Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Xin, Nie, Ping, Guo, Yiwen, Wei, Haojie, Zhang, Zhanqiu, Minervini, Pasquale, Ma, Ruotian, Gui, Tao, Zhang, Qi, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Cultivating Game Sense for Yourself: Making VLMs Gaming Experts
von: Lu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenxuan, et al.
Veröffentlicht: (2025)
Answerability in Retrieval-Augmented Open-Domain Question Answering
von: Abdumalikov, Rustam, et al.
Veröffentlicht: (2024)
von: Abdumalikov, Rustam, et al.
Veröffentlicht: (2024)
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
von: Li, Muqing, et al.
Veröffentlicht: (2025)
von: Li, Muqing, et al.
Veröffentlicht: (2025)
Are Large Language Models Good Prompt Optimizers?
von: Ma, Ruotian, et al.
Veröffentlicht: (2024)
von: Ma, Ruotian, et al.
Veröffentlicht: (2024)
Unveiling Linguistic Regions in Large Language Models
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
ContextWIN: Whittle Index Based Mixture-of-Experts Neural Model For Restless Bandits Via Deep RL
von: Guo, Zhanqiu, et al.
Veröffentlicht: (2024)
von: Guo, Zhanqiu, et al.
Veröffentlicht: (2024)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather Forecasting
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
von: Ye, Xin, et al.
Veröffentlicht: (2026)
von: Ye, Xin, et al.
Veröffentlicht: (2026)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
Do Domain-specific Experts exist in MoE-based LLMs?
von: Do, Giang, et al.
Veröffentlicht: (2026)
von: Do, Giang, et al.
Veröffentlicht: (2026)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
Advancing Expert Specialization for Better MoE
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
von: Zhao, Ziyu, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2026)
PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs
von: Zhang, Ze Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Ze Yu, et al.
Veröffentlicht: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
von: Yun, Sukwon, et al.
Veröffentlicht: (2024)
von: Yun, Sukwon, et al.
Veröffentlicht: (2024)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
von: Zhou, Xin, et al.
Veröffentlicht: (2025) -
Cultivating Game Sense for Yourself: Making VLMs Gaming Experts
von: Lu, Wenxuan, et al.
Veröffentlicht: (2025) -
Answerability in Retrieval-Augmented Open-Domain Question Answering
von: Abdumalikov, Rustam, et al.
Veröffentlicht: (2024) -
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
von: Li, Muqing, et al.
Veröffentlicht: (2025) -
Are Large Language Models Good Prompt Optimizers?
von: Ma, Ruotian, et al.
Veröffentlicht: (2024)