Improving the Downstream Performance of Mixture-of-Experts Transformers via Weak Vanilla Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Xin, Zhao, Yanyan, Qin, Bing, Liu, Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
von: Lu, Xin, et al.
Veröffentlicht: (2024)
von: Lu, Xin, et al.
Veröffentlicht: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025)
von: Lu, Xin, et al.
Veröffentlicht: (2025)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset
von: Le, Thanh-Dung, et al.
Veröffentlicht: (2023)
von: Le, Thanh-Dung, et al.
Veröffentlicht: (2023)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
Superposition in Transformers: A Novel Way of Building Mixture of Experts
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
von: Jung, Min Jae, et al.
Veröffentlicht: (2024)
von: Jung, Min Jae, et al.
Veröffentlicht: (2024)
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
von: Zhao, Liang, et al.
Veröffentlicht: (2023)
von: Zhao, Liang, et al.
Veröffentlicht: (2023)
Mixture of Hidden-Dimensions Transformer
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Enhancing Complex Causality Extraction via Improved Subtask Interaction and Knowledge Fusion
von: Gao, Jinglong, et al.
Veröffentlicht: (2024)
von: Gao, Jinglong, et al.
Veröffentlicht: (2024)
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
Data Uncertainty-Aware Learning for Multimodal Aspect-based Sentiment Analysis
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
von: Su, Guinan, et al.
Veröffentlicht: (2025)
von: Su, Guinan, et al.
Veröffentlicht: (2025)
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
von: Wang, Bichen, et al.
Veröffentlicht: (2025)
von: Wang, Bichen, et al.
Veröffentlicht: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Task-Routed Mixture-of-Experts with Cognitive Appraisal for Implicit Sentiment Analysis
von: Chai, Yaping, et al.
Veröffentlicht: (2026)
von: Chai, Yaping, et al.
Veröffentlicht: (2026)
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts
von: Liu, Weisi, et al.
Veröffentlicht: (2025)
von: Liu, Weisi, et al.
Veröffentlicht: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
StockBot 2.0: Vanilla LSTMs Outperform Transformer-based Forecasting for Stock Prices
von: Mohanty, Shaswat
Veröffentlicht: (2026)
von: Mohanty, Shaswat
Veröffentlicht: (2026)
Mixture of Neuron Experts
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
RKLD: Reverse KL-Divergence-based Knowledge Distillation for Unlearning Personal Information in Large Language Models
von: Wang, Bichen, et al.
Veröffentlicht: (2024)
von: Wang, Bichen, et al.
Veröffentlicht: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
von: Lu, Xin, et al.
Veröffentlicht: (2024) -
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025) -
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
von: Wu, Di, et al.
Veröffentlicht: (2026) -
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
von: Wu, Di, et al.
Veröffentlicht: (2024) -
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset
von: Le, Thanh-Dung, et al.
Veröffentlicht: (2023)