Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Liu, Zhiyuan, Li, Yueyu, Lin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
State Tuning: State-based Test-Time Scaling on RWKV-7
por: Xiao, Liu, et al.
Publicado: (2025)
por: Xiao, Liu, et al.
Publicado: (2025)
Cross-attention for State-based model RWKV-7
por: Xiao, Liu, et al.
Publicado: (2025)
por: Xiao, Liu, et al.
Publicado: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
por: Xia, Xinfeng, et al.
Publicado: (2025)
por: Xia, Xinfeng, et al.
Publicado: (2025)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
por: Lee, Jaeseong, et al.
Publicado: (2024)
por: Lee, Jaeseong, et al.
Publicado: (2024)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
por: Yuan, Yueming, et al.
Publicado: (2025)
por: Yuan, Yueming, et al.
Publicado: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
por: Gu, Naibin, et al.
Publicado: (2025)
por: Gu, Naibin, et al.
Publicado: (2025)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
por: Manzoni, Andrea
Publicado: (2026)
por: Manzoni, Andrea
Publicado: (2026)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
por: Li, Shuhuai, et al.
Publicado: (2026)
por: Li, Shuhuai, et al.
Publicado: (2026)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
por: Zeng, Zhiyuan, et al.
Publicado: (2024)
por: Zeng, Zhiyuan, et al.
Publicado: (2024)
GRIN: GRadient-INformed MoE
por: Liu, Liyuan, et al.
Publicado: (2024)
por: Liu, Liyuan, et al.
Publicado: (2024)
WuNeng: Hybrid State with Attention
por: Xiao, Liu, et al.
Publicado: (2025)
por: Xiao, Liu, et al.
Publicado: (2025)
Faster MoE LLM Inference for Extremely Large Models
por: Yang, Haoqi, et al.
Publicado: (2025)
por: Yang, Haoqi, et al.
Publicado: (2025)
LocMoE: A Low-Overhead MoE for Large Language Model Training
por: Li, Jing, et al.
Publicado: (2024)
por: Li, Jing, et al.
Publicado: (2024)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
por: Li, Zichong, et al.
Publicado: (2025)
por: Li, Zichong, et al.
Publicado: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
por: Cao, Mingyu, et al.
Publicado: (2024)
por: Cao, Mingyu, et al.
Publicado: (2024)
Steering MoE LLMs via Expert (De)Activation
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
por: Deng, Jianing, et al.
Publicado: (2026)
por: Deng, Jianing, et al.
Publicado: (2026)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
por: Antoniak, Szymon, et al.
Publicado: (2023)
por: Antoniak, Szymon, et al.
Publicado: (2023)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
por: Sun, Libo, et al.
Publicado: (2026)
por: Sun, Libo, et al.
Publicado: (2026)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
por: Peng, Bo, et al.
Publicado: (2025)
por: Peng, Bo, et al.
Publicado: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
por: Huang, Quzhe, et al.
Publicado: (2024)
por: Huang, Quzhe, et al.
Publicado: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024)
por: Muzio, Alexandre, et al.
Publicado: (2024)
Towards an empirical understanding of MoE design choices
por: Fan, Dongyang, et al.
Publicado: (2024)
por: Fan, Dongyang, et al.
Publicado: (2024)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
por: Liu, Xiangyue, et al.
Publicado: (2026)
por: Liu, Xiangyue, et al.
Publicado: (2026)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
por: Shi, Chufan, et al.
Publicado: (2024)
por: Shi, Chufan, et al.
Publicado: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
por: Kang, Junmo, et al.
Publicado: (2024)
por: Kang, Junmo, et al.
Publicado: (2024)
ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
por: Zhao, Ziyu, et al.
Publicado: (2026)
por: Zhao, Ziyu, et al.
Publicado: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
por: Sun, Yitong, et al.
Publicado: (2026)
por: Sun, Yitong, et al.
Publicado: (2026)
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
por: Xu, Yuzhuang, et al.
Publicado: (2025)
por: Xu, Yuzhuang, et al.
Publicado: (2025)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
por: Zhou, Wenyong, et al.
Publicado: (2026)
por: Zhou, Wenyong, et al.
Publicado: (2026)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
por: Wang, Bo, et al.
Publicado: (2026)
por: Wang, Bo, et al.
Publicado: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
por: Chen, Guanjie, et al.
Publicado: (2024)
por: Chen, Guanjie, et al.
Publicado: (2024)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
por: Li, Pingzhi, et al.
Publicado: (2025)
por: Li, Pingzhi, et al.
Publicado: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
por: Ma, Wenhan, et al.
Publicado: (2025)
por: Ma, Wenhan, et al.
Publicado: (2025)
Ejemplares similares
-
State Tuning: State-based Test-Time Scaling on RWKV-7
por: Xiao, Liu, et al.
Publicado: (2025) -
Cross-attention for State-based model RWKV-7
por: Xiao, Liu, et al.
Publicado: (2025) -
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
por: Xia, Xinfeng, et al.
Publicado: (2025) -
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
por: Lee, Jaeseong, et al.
Publicado: (2024) -
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
por: Yuan, Yueming, et al.
Publicado: (2025)