BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Qizhen, Gritsch, Nikolas, Gnaneshwar, Dwaraknath, Guo, Simon, Cairuz, David, Venkitesh, Bharat, Foerster, Jakob, Blunsom, Phil, Ruder, Sebastian, Ustun, Ahmet, Locatelli, Acyr |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
Rope to Nope and Back Again: A New Hybrid Attention Strategy
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
Aya 23: Open Weight Releases to Further Multilingual Progress
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
SnapKV: LLM Knows What You are Looking for Before Generation
von: Li, Yuhong, et al.
Veröffentlicht: (2024)
von: Li, Yuhong, et al.
Veröffentlicht: (2024)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
PARDEN, Can You Repeat That? Defending against Jailbreaks via Repetition
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
Human Feedback is not Gold Standard
von: Hosking, Tom, et al.
Veröffentlicht: (2023)
von: Hosking, Tom, et al.
Veröffentlicht: (2023)
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2025)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2025)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
von: Abagyan, Diana, et al.
Veröffentlicht: (2025)
von: Abagyan, Diana, et al.
Veröffentlicht: (2025)
Polynomials, Divided Differences, and Codes
von: Venkitesh, S.
Veröffentlicht: (2024)
von: Venkitesh, S.
Veröffentlicht: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
How Does Quantization Affect Multilingual LLMs?
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2026)
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2024)
von: Fu, Li, et al.
Veröffentlicht: (2024)
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
von: Huang, Yongqi, et al.
Veröffentlicht: (2025)
von: Huang, Yongqi, et al.
Veröffentlicht: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
O que os analistas pensam sobre a homossexualidade?
von: Acyr Maya
Veröffentlicht: (2007)
von: Acyr Maya
Veröffentlicht: (2007)
An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence
von: Zhang, Qizhen, et al.
Veröffentlicht: (2026)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2026)
Analysing the Sample Complexity of Opponent Shaping
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
Random Reed-Solomon Codes are List Recoverable with Optimal List Size
von: Doron, Dean, et al.
Veröffentlicht: (2024)
von: Doron, Dean, et al.
Veröffentlicht: (2024)
List Recoverable Codes: The Good, the Bad, and the Unknown (hopefully not Ugly)
von: Resch, Nicolas, et al.
Veröffentlicht: (2025)
von: Resch, Nicolas, et al.
Veröffentlicht: (2025)
Optical single-shot readout of spin qubits in silicon
von: Gritsch, Andreas, et al.
Veröffentlicht: (2024)
von: Gritsch, Andreas, et al.
Veröffentlicht: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024) -
Rope to Nope and Back Again: A New Hybrid Attention Strategy
von: Yang, Bowen, et al.
Veröffentlicht: (2025) -
Aya 23: Open Weight Releases to Further Multilingual Progress
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024) -
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026) -
SnapKV: LLM Knows What You are Looking for Before Generation
von: Li, Yuhong, et al.
Veröffentlicht: (2024)