FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Juyong, Wang, Fan, Qi, Hong, Kim, Sunghun, Tang, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
di: Cai, Weilin, et al.
Pubblicazione: (2024)
di: Cai, Weilin, et al.
Pubblicazione: (2024)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
di: Xue, Fuzhao, et al.
Pubblicazione: (2024)
di: Xue, Fuzhao, et al.
Pubblicazione: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
di: Wawdhane, Sourish, et al.
Pubblicazione: (2026)
di: Wawdhane, Sourish, et al.
Pubblicazione: (2026)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
di: Singh, Gursimran, et al.
Pubblicazione: (2025)
di: Singh, Gursimran, et al.
Pubblicazione: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
di: Qian, Yulei, et al.
Pubblicazione: (2024)
di: Qian, Yulei, et al.
Pubblicazione: (2024)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
di: Yuan, Yueming, et al.
Pubblicazione: (2025)
di: Yuan, Yueming, et al.
Pubblicazione: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
di: Luo, Shuqing, et al.
Pubblicazione: (2024)
di: Luo, Shuqing, et al.
Pubblicazione: (2024)
BlackMamba: Mixture of Experts for State-Space Models
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs
di: Park, Chansung, et al.
Pubblicazione: (2024)
di: Park, Chansung, et al.
Pubblicazione: (2024)
Scalable Training of Mixture-of-Experts Models with Megatron Core
di: Yan, Zijie, et al.
Pubblicazione: (2026)
di: Yan, Zijie, et al.
Pubblicazione: (2026)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
di: Jin, Chao, et al.
Pubblicazione: (2025)
di: Jin, Chao, et al.
Pubblicazione: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
di: Gao, Yunqi, et al.
Pubblicazione: (2025)
di: Gao, Yunqi, et al.
Pubblicazione: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
di: Liu, Ziming, et al.
Pubblicazione: (2025)
di: Liu, Ziming, et al.
Pubblicazione: (2025)
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
di: Wang, Xinbo, et al.
Pubblicazione: (2025)
di: Wang, Xinbo, et al.
Pubblicazione: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
di: Shi, Long, et al.
Pubblicazione: (2025)
di: Shi, Long, et al.
Pubblicazione: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
di: Yang, Zheming, et al.
Pubblicazione: (2025)
di: Yang, Zheming, et al.
Pubblicazione: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
di: Vooturi, Dharma Teja, et al.
Pubblicazione: (2026)
di: Vooturi, Dharma Teja, et al.
Pubblicazione: (2026)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
di: Du, Zhixu, et al.
Pubblicazione: (2023)
di: Du, Zhixu, et al.
Pubblicazione: (2023)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
di: Tang, Xinru, et al.
Pubblicazione: (2025)
di: Tang, Xinru, et al.
Pubblicazione: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
di: Wang, Wenfeng, et al.
Pubblicazione: (2025)
di: Wang, Wenfeng, et al.
Pubblicazione: (2025)
FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models
di: Zhang, Zishuai, et al.
Pubblicazione: (2025)
di: Zhang, Zishuai, et al.
Pubblicazione: (2025)
CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models
di: Jiang, Lei, et al.
Pubblicazione: (2025)
di: Jiang, Lei, et al.
Pubblicazione: (2025)
Inference Acceleration for Large Language Models on CPUs
di: PS, Ditto, et al.
Pubblicazione: (2024)
di: PS, Ditto, et al.
Pubblicazione: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
di: Wu, Yongji, et al.
Pubblicazione: (2025)
di: Wu, Yongji, et al.
Pubblicazione: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
di: Tairin, Suraiya, et al.
Pubblicazione: (2025)
di: Tairin, Suraiya, et al.
Pubblicazione: (2025)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
di: Zhang, Mohan, et al.
Pubblicazione: (2025)
di: Zhang, Mohan, et al.
Pubblicazione: (2025)
A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine
di: Li, Anran, et al.
Pubblicazione: (2026)
di: Li, Anran, et al.
Pubblicazione: (2026)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
di: Brown, Nick, et al.
Pubblicazione: (2025)
di: Brown, Nick, et al.
Pubblicazione: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
di: Xue, Nan, et al.
Pubblicazione: (2024)
di: Xue, Nan, et al.
Pubblicazione: (2024)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)
di: Shen, Zixu, et al.
Pubblicazione: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
di: Wang, Haodong, et al.
Pubblicazione: (2025)
di: Wang, Haodong, et al.
Pubblicazione: (2025)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
di: Huang, You-Liang, et al.
Pubblicazione: (2026)
di: Huang, You-Liang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
di: Cai, Weilin, et al.
Pubblicazione: (2024) -
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
di: Xue, Fuzhao, et al.
Pubblicazione: (2024) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
di: Sun, Weigao, et al.
Pubblicazione: (2025) -
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
di: Wawdhane, Sourish, et al.
Pubblicazione: (2026) -
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
di: Singh, Gursimran, et al.
Pubblicazione: (2025)