UMoE: Unifying Attention and FFN with Shared Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Yuanhang, Wang, Chaozheng, Li, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
por: Liu, Ningyuan, et al.
Publicado: (2025)
por: Liu, Ningyuan, et al.
Publicado: (2025)
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
por: Song, Chendong, et al.
Publicado: (2026)
por: Song, Chendong, et al.
Publicado: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
por: Pei, Zehua, et al.
Publicado: (2025)
por: Pei, Zehua, et al.
Publicado: (2025)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
por: Smithline, Gabriel, et al.
Publicado: (2026)
por: Smithline, Gabriel, et al.
Publicado: (2026)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
por: Gautam, Aayush, et al.
Publicado: (2026)
por: Gautam, Aayush, et al.
Publicado: (2026)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
por: Jing, Phoebe, et al.
Publicado: (2024)
por: Jing, Phoebe, et al.
Publicado: (2024)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
por: Jaiswal, Ajay, et al.
Publicado: (2025)
por: Jaiswal, Ajay, et al.
Publicado: (2025)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
por: Zhao, Yanjun, et al.
Publicado: (2024)
por: Zhao, Yanjun, et al.
Publicado: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
por: Huang, Minbin, et al.
Publicado: (2026)
por: Huang, Minbin, et al.
Publicado: (2026)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
por: Wu, Hanjiang, et al.
Publicado: (2026)
por: Wu, Hanjiang, et al.
Publicado: (2026)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
por: Liu, Renpu, et al.
Publicado: (2025)
por: Liu, Renpu, et al.
Publicado: (2025)
Quaternion Self-Attention with Shared Scores
por: Yamauchi, Shogo, et al.
Publicado: (2026)
por: Yamauchi, Shogo, et al.
Publicado: (2026)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
por: Li, Junzhuo, et al.
Publicado: (2026)
por: Li, Junzhuo, et al.
Publicado: (2026)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
por: Li, Wenbing, et al.
Publicado: (2025)
por: Li, Wenbing, et al.
Publicado: (2025)
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
por: Yang, Yuanhang, et al.
Publicado: (2024)
por: Yang, Yuanhang, et al.
Publicado: (2024)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
por: Pan, Yu, et al.
Publicado: (2025)
por: Pan, Yu, et al.
Publicado: (2025)
UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
por: Wang, Fu-Yun, et al.
Publicado: (2025)
por: Wang, Fu-Yun, et al.
Publicado: (2025)
Adaptive Shared Experts with LoRA-Based Mixture of Experts for Multi-Task Learning
por: Yang, Minghao, et al.
Publicado: (2025)
por: Yang, Minghao, et al.
Publicado: (2025)
Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
por: Yuan, Li, et al.
Publicado: (2025)
por: Yuan, Li, et al.
Publicado: (2025)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
por: Chou, Yuhong, et al.
Publicado: (2024)
por: Chou, Yuhong, et al.
Publicado: (2024)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
por: Kunwar, Pradip, et al.
Publicado: (2025)
por: Kunwar, Pradip, et al.
Publicado: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
por: Fu, Zichuan, et al.
Publicado: (2026)
por: Fu, Zichuan, et al.
Publicado: (2026)
MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction
por: Wang, Xiaoyang, et al.
Publicado: (2025)
por: Wang, Xiaoyang, et al.
Publicado: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
por: Yang, Cheng, et al.
Publicado: (2024)
por: Yang, Cheng, et al.
Publicado: (2024)
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
por: Zhao, Ziyu, et al.
Publicado: (2025)
por: Zhao, Ziyu, et al.
Publicado: (2025)
Unified Class and Domain Incremental Learning with Mixture of Experts for Indoor Localization
por: Singampalli, Akhil, et al.
Publicado: (2025)
por: Singampalli, Akhil, et al.
Publicado: (2025)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
por: Vankov, Daniil, et al.
Publicado: (2026)
por: Vankov, Daniil, et al.
Publicado: (2026)
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
por: Pati, Viresh, et al.
Publicado: (2026)
por: Pati, Viresh, et al.
Publicado: (2026)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
por: Gao, Yifei, et al.
Publicado: (2025)
por: Gao, Yifei, et al.
Publicado: (2025)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
por: Hao, Jiawei, et al.
Publicado: (2026)
por: Hao, Jiawei, et al.
Publicado: (2026)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
por: Li, Xinyi, et al.
Publicado: (2025)
por: Li, Xinyi, et al.
Publicado: (2025)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
por: Chen, Yuanteng, et al.
Publicado: (2025)
por: Chen, Yuanteng, et al.
Publicado: (2025)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
por: Huang, Ruijun, et al.
Publicado: (2026)
por: Huang, Ruijun, et al.
Publicado: (2026)
Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance
por: Li, Jihang, et al.
Publicado: (2026)
por: Li, Jihang, et al.
Publicado: (2026)
HiF-DTA: Hierarchical Feature Learning Network for Drug-Target Affinity Prediction
por: Li, Minghui, et al.
Publicado: (2025)
por: Li, Minghui, et al.
Publicado: (2025)
Ejemplares similares
-
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
por: Liu, Ningyuan, et al.
Publicado: (2025) -
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
por: Song, Chendong, et al.
Publicado: (2026) -
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
por: Li, Cheng, et al.
Publicado: (2025) -
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
por: Pei, Zehua, et al.
Publicado: (2025) -
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
por: Smithline, Gabriel, et al.
Publicado: (2026)