Advancing Expert Specialization for Better MoE
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Hongcan, Lu, Haolang, Nan, Guoshun, Chu, Bolun, Zhuang, Jialin, Yang, Yuan, Che, Wenhao, Cao, Xinye, Leng, Sicong, Cui, Qimei, Jiang, Xudong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
by: Lu, Haolang, et al.
Published: (2025)
by: Lu, Haolang, et al.
Published: (2025)
KGMark: A Diffusion Watermark for Knowledge Graphs
by: Peng, Hongrui, et al.
Published: (2025)
by: Peng, Hongrui, et al.
Published: (2025)
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
by: Gu, Jintao, et al.
Published: (2025)
by: Gu, Jintao, et al.
Published: (2025)
Disentangling Deception and Hallucination Failures in LLMs
by: Lu, Haolang, et al.
Published: (2026)
by: Lu, Haolang, et al.
Published: (2026)
Negative Curves and Elliptic Fibrations on a Special Rational Surface
by: Mendes, Luís Gustavo, et al.
Published: (2024)
by: Mendes, Luís Gustavo, et al.
Published: (2024)
A proof of the uniqueness of the limit cycle of a quasi-homogeneous system
by: Zhuang, Ziwei, et al.
Published: (2022)
by: Zhuang, Ziwei, et al.
Published: (2022)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)
by: Vankov, Daniil, et al.
Published: (2026)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
by: Feng, Ruitao, et al.
Published: (2025)
by: Feng, Ruitao, et al.
Published: (2025)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
by: Lasby, Mike, et al.
Published: (2025)
by: Lasby, Mike, et al.
Published: (2025)
An efficient search strategy for hidden ideals in pointed partially ordered sets
by: Eisel, Roma, et al.
Published: (2025)
by: Eisel, Roma, et al.
Published: (2025)
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026)
by: Rai, Daking, et al.
Published: (2026)
Scalify: scale propagation for efficient low-precision LLM training
by: Balança, Paul, et al.
Published: (2024)
by: Balança, Paul, et al.
Published: (2024)
Extending $μ$P: Spectral Conditions for Feature Learning Across Optimizers
by: Gupta, Akshita, et al.
Published: (2026)
by: Gupta, Akshita, et al.
Published: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
by: Mathew, Aby Mammen
Published: (2026)
by: Mathew, Aby Mammen
Published: (2026)
The moduli space of Hermitian-Yang-Mills connections
by: Sasaki, Jun
Published: (2025)
by: Sasaki, Jun
Published: (2025)
The moduli space of Higgs pairs
by: Sasaki, Jun
Published: (2026)
by: Sasaki, Jun
Published: (2026)
Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Quadratic Forms, Exact Covering Systems, and Product Identities for Theta Functions
by: Cao, Zhu
Published: (2025)
by: Cao, Zhu
Published: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
by: Cao, Xinye, et al.
Published: (2025)
by: Cao, Xinye, et al.
Published: (2025)
Continuous Latent Diffusion Language Model
by: Guo, Hongcan, et al.
Published: (2026)
by: Guo, Hongcan, et al.
Published: (2026)
Best-fit Weibull distributions of sediment core OC437-07_GC27
by: McGee, David, et al.
Published: (2013)
by: McGee, David, et al.
Published: (2013)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
Supplementary Materials to Graph Convolutional Branch and Bound
by: Sciandra, Lorenzo, et al.
Published: (2024)
by: Sciandra, Lorenzo, et al.
Published: (2024)
Few-shot crack image classification using clip based on bayesian optimization
by: Zhang, Yingchao, et al.
Published: (2025)
by: Zhang, Yingchao, et al.
Published: (2025)
Probing for Representation Manifolds in Superposition
by: Modell, Alexander
Published: (2026)
by: Modell, Alexander
Published: (2026)
The Origins of Representation Manifolds in Large Language Models
by: Modell, Alexander, et al.
Published: (2025)
by: Modell, Alexander, et al.
Published: (2025)
HR-Agent: A Task-Oriented Dialogue (TOD) LLM Agent Tailored for HR Applications
by: Xu, Weijie, et al.
Published: (2024)
by: Xu, Weijie, et al.
Published: (2024)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025)
by: Danieli, Federico, et al.
Published: (2025)
On Krull dimension of modules over group rings of minimax abelian groups
by: Tushev, Anatolii V.
Published: (2024)
by: Tushev, Anatolii V.
Published: (2024)
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025)
by: Chahine, Makram, et al.
Published: (2025)
Non-Asymptotic Convergence of Discrete Diffusion Models: Masked and Random Walk dynamics
by: Conforti, Giovanni, et al.
Published: (2025)
by: Conforti, Giovanni, et al.
Published: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
by: Tang, Zhengzheng
Published: (2026)
by: Tang, Zhengzheng
Published: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
by: Harvey, Daniel Fidel, et al.
Published: (2025)
by: Harvey, Daniel Fidel, et al.
Published: (2025)
Similar Items
-
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
by: Lu, Haolang, et al.
Published: (2025) -
KGMark: A Diffusion Watermark for Knowledge Graphs
by: Peng, Hongrui, et al.
Published: (2025) -
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025) -
Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
by: Gu, Jintao, et al.
Published: (2025) -
Disentangling Deception and Hallucination Failures in LLMs
by: Lu, Haolang, et al.
Published: (2026)