DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Chenyang, Zhao, Weilin, Han, Xu, Xiao, Chaojun, Chen, Yingfa, Liu, Zhiyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
di: Song, Chenyang, et al.
Pubblicazione: (2025)
di: Song, Chenyang, et al.
Pubblicazione: (2025)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
di: Luo, Yuqi, et al.
Pubblicazione: (2024)
di: Luo, Yuqi, et al.
Pubblicazione: (2024)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
Robust and Scalable Model Editing for Large Language Models
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
NOSA: Native and Offloadable Sparse Attention
di: Huang, Yuxiang, et al.
Pubblicazione: (2025)
di: Huang, Yuxiang, et al.
Pubblicazione: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
di: Chen, Yingfa, et al.
Pubblicazione: (2025)
di: Chen, Yingfa, et al.
Pubblicazione: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
di: Pan, Bowen, et al.
Pubblicazione: (2024)
di: Pan, Bowen, et al.
Pubblicazione: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
di: Huang, Yuxiang, et al.
Pubblicazione: (2025)
di: Huang, Yuxiang, et al.
Pubblicazione: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
Configurable Foundation Models: Building LLMs from a Modular Perspective
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
A Survey on Mixture of Experts in Large Language Models
di: Cai, Weilin, et al.
Pubblicazione: (2024)
di: Cai, Weilin, et al.
Pubblicazione: (2024)
47B Mixture-of-Experts Beats 671B Dense Models on Chinese Medical Examinations
di: Tseng, Chiung-Yi, et al.
Pubblicazione: (2025)
di: Tseng, Chiung-Yi, et al.
Pubblicazione: (2025)
StateX: Enhancing RNN Recall via Post-training State Expansion
di: Shen, Xingyu, et al.
Pubblicazione: (2025)
di: Shen, Xingyu, et al.
Pubblicazione: (2025)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
di: Chen, Guanjie, et al.
Pubblicazione: (2024)
di: Chen, Guanjie, et al.
Pubblicazione: (2024)
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering
di: Verma, Vinay Kumar, et al.
Pubblicazione: (2025)
di: Verma, Vinay Kumar, et al.
Pubblicazione: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
di: He, Shwai, et al.
Pubblicazione: (2025)
di: He, Shwai, et al.
Pubblicazione: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
di: Chen, Yanbei, et al.
Pubblicazione: (2026)
di: Chen, Yanbei, et al.
Pubblicazione: (2026)
CoSMoEs: Compact Sparse Mixture of Experts
di: Huber, Patrick, et al.
Pubblicazione: (2025)
di: Huber, Patrick, et al.
Pubblicazione: (2025)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
di: Ai, Mengting, et al.
Pubblicazione: (2023)
di: Ai, Mengting, et al.
Pubblicazione: (2023)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
di: Muzio, Alexandre, et al.
Pubblicazione: (2024)
di: Muzio, Alexandre, et al.
Pubblicazione: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
di: Cai, Weilin, et al.
Pubblicazione: (2024)
di: Cai, Weilin, et al.
Pubblicazione: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
di: Dong, Zican, et al.
Pubblicazione: (2025)
di: Dong, Zican, et al.
Pubblicazione: (2025)
Mixture of Lookup Experts
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
di: Kang, Hao, et al.
Pubblicazione: (2025)
di: Kang, Hao, et al.
Pubblicazione: (2025)
Routing-Free Mixture-of-Experts
di: Liu, Yilun, et al.
Pubblicazione: (2026)
di: Liu, Yilun, et al.
Pubblicazione: (2026)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
di: Gao, Shangqian, et al.
Pubblicazione: (2025)
di: Gao, Shangqian, et al.
Pubblicazione: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
di: Zhang, Zhengyan, et al.
Pubblicazione: (2024)
di: Zhang, Zhengyan, et al.
Pubblicazione: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
di: Wang, An, et al.
Pubblicazione: (2024)
di: Wang, An, et al.
Pubblicazione: (2024)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
di: Song, Jonghyun, et al.
Pubblicazione: (2025)
di: Song, Jonghyun, et al.
Pubblicazione: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
di: Wang, Lean, et al.
Pubblicazione: (2024)
di: Wang, Lean, et al.
Pubblicazione: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
di: Song, Guanghui, et al.
Pubblicazione: (2025)
di: Song, Guanghui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
di: Song, Chenyang, et al.
Pubblicazione: (2025) -
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
di: Luo, Yuqi, et al.
Pubblicazione: (2024) -
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
di: Zhao, Weilin, et al.
Pubblicazione: (2025) -
Robust and Scalable Model Editing for Large Language Models
di: Chen, Yingfa, et al.
Pubblicazione: (2024) -
NOSA: Native and Offloadable Sparse Attention
di: Huang, Yuxiang, et al.
Pubblicazione: (2025)