Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chung, Youngseog, Malik, Dhruv, Schneider, Jeff, Li, Yuanzhi, Singh, Aarti |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025)
von: Liu, Grace, et al.
Veröffentlicht: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
von: Rastegar, Reza
Veröffentlicht: (2026)
von: Rastegar, Reza
Veröffentlicht: (2026)
Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNs
von: Lahoti, Aakash, et al.
Veröffentlicht: (2024)
von: Lahoti, Aakash, et al.
Veröffentlicht: (2024)
Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts
von: Ubukata, Toshihide, et al.
Veröffentlicht: (2026)
von: Ubukata, Toshihide, et al.
Veröffentlicht: (2026)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
von: Zou, Heming, et al.
Veröffentlicht: (2025)
von: Zou, Heming, et al.
Veröffentlicht: (2025)
Mixture of Concept Bottleneck Experts
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
Specifying and Solving Robust Empirical Risk Minimization Problems Using CVXPY
von: Luxenberg, Eric, et al.
Veröffentlicht: (2023)
von: Luxenberg, Eric, et al.
Veröffentlicht: (2023)
From Sparse to Soft Mixtures of Experts
von: Puigcerver, Joan, et al.
Veröffentlicht: (2023)
von: Puigcerver, Joan, et al.
Veröffentlicht: (2023)
Cooperative Multi-agent RL with Communication Constraints
von: Xiong, Nuoya, et al.
Veröffentlicht: (2026)
von: Xiong, Nuoya, et al.
Veröffentlicht: (2026)
Projection Optimization: A General Framework for Multi-Objective and Multi-Group RLHF
von: Xiong, Nuoya, et al.
Veröffentlicht: (2025)
von: Xiong, Nuoya, et al.
Veröffentlicht: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
von: Elango, Venmugil, et al.
Veröffentlicht: (2026)
von: Elango, Venmugil, et al.
Veröffentlicht: (2026)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
von: Wang, Guorun, et al.
Veröffentlicht: (2024)
von: Wang, Guorun, et al.
Veröffentlicht: (2024)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Theory on Mixture-of-Experts in Continual Learning
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
von: Huang, Yu, et al.
Veröffentlicht: (2026)
von: Huang, Yu, et al.
Veröffentlicht: (2026)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Mixture of Diverse Size Experts
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
Mixture of A Million Experts
von: He, Xu Owen
Veröffentlicht: (2024)
von: He, Xu Owen
Veröffentlicht: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Sparsity and Superposition in Mixture of Experts
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
von: Liu, Ningyuan, et al.
Veröffentlicht: (2025)
von: Liu, Ningyuan, et al.
Veröffentlicht: (2025)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
von: Kunwar, Pradip, et al.
Veröffentlicht: (2025)
von: Kunwar, Pradip, et al.
Veröffentlicht: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Mixture-of-Experts Meets In-Context Reinforcement Learning
von: Wu, Wenhao, et al.
Veröffentlicht: (2025)
von: Wu, Wenhao, et al.
Veröffentlicht: (2025)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
von: Liu, Feilong
Veröffentlicht: (2026)
von: Liu, Feilong
Veröffentlicht: (2026)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning
von: Feng, Yu, et al.
Veröffentlicht: (2025)
von: Feng, Yu, et al.
Veröffentlicht: (2025)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025) -
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024) -
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024) -
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
von: Rastegar, Reza
Veröffentlicht: (2026) -
Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNs
von: Lahoti, Aakash, et al.
Veröffentlicht: (2024)