XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Yuanhang, Qi, Shiyi, Gu, Wenchao, Wang, Chaozheng, Gao, Cuiyun, Xu, Zenglin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UMoE: Unifying Attention and FFN with Shared Experts
von: Yang, Yuanhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuanhang, et al.
Veröffentlicht: (2025)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
Enhancing Multivariate Time Series Forecasting with Mutual Information-driven Cross-Variable and Temporal Modeling
von: Qi, Shiyi, et al.
Veröffentlicht: (2024)
von: Qi, Shiyi, et al.
Veröffentlicht: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
von: Jing, Phoebe, et al.
Veröffentlicht: (2024)
von: Jing, Phoebe, et al.
Veröffentlicht: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
Revisiting Long-term Time Series Forecasting: An Investigation on Linear Mapping
von: Li, Zhe, et al.
Veröffentlicht: (2023)
von: Li, Zhe, et al.
Veröffentlicht: (2023)
Adaptive and Fine-grained Module-wise Expert Pruning for Efficient LoRA-MoE Fine-Tuning
von: Li, Weihang, et al.
Veröffentlicht: (2026)
von: Li, Weihang, et al.
Veröffentlicht: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
Aligning Large Language Models via Fine-grained Supervision
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
von: Zhang, Shuibai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuibai, et al.
Veröffentlicht: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
von: Wang, Zige, et al.
Veröffentlicht: (2025)
von: Wang, Zige, et al.
Veröffentlicht: (2025)
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
Adaptive LoRA Experts Allocation and Selection for Federated Fine-Tuning
von: Wang, Lei, et al.
Veröffentlicht: (2025)
von: Wang, Lei, et al.
Veröffentlicht: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
von: Wang, Qibin, et al.
Veröffentlicht: (2024)
von: Wang, Qibin, et al.
Veröffentlicht: (2024)
CoSMoEs: Compact Sparse Mixture of Experts
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
von: Liu, Zhihao, et al.
Veröffentlicht: (2024)
von: Liu, Zhihao, et al.
Veröffentlicht: (2024)
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UMoE: Unifying Attention and FFN with Shared Experts
von: Yang, Yuanhang, et al.
Veröffentlicht: (2025) -
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
von: Wang, Xinyi, et al.
Veröffentlicht: (2025) -
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2024) -
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
von: Tan, Zheyue, et al.
Veröffentlicht: (2025) -
Enhancing Multivariate Time Series Forecasting with Mutual Information-driven Cross-Variable and Temporal Modeling
von: Qi, Shiyi, et al.
Veröffentlicht: (2024)