Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Boan, Ding, Liang, Shen, Li, Peng, Keqin, Cao, Yu, Cheng, Dazhao, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
par: Xu, Ziyang, et autres
Publié: (2024)
par: Xu, Ziyang, et autres
Publié: (2024)
Revisiting Catastrophic Forgetting in Large Language Model Tuning
par: Li, Hongyu, et autres
Publié: (2024)
par: Li, Hongyu, et autres
Publié: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
par: Zan, Changtong, et autres
Publié: (2024)
par: Zan, Changtong, et autres
Publié: (2024)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
par: Lu, Qingyu, et autres
Publié: (2025)
par: Lu, Qingyu, et autres
Publié: (2025)
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
par: Wu, Yuchen, et autres
Publié: (2025)
par: Wu, Yuchen, et autres
Publié: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)
par: Ma, Zhicheng, et autres
Publié: (2026)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
par: Wei, Tianwen, et autres
Publié: (2024)
par: Wei, Tianwen, et autres
Publié: (2024)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
par: Sun, Liangtai, et autres
Publié: (2024)
par: Sun, Liangtai, et autres
Publié: (2024)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
par: Zhu, Shien, et autres
Publié: (2025)
par: Zhu, Shien, et autres
Publié: (2025)
Upcycling Large Language Models into Mixture of Experts
par: He, Ethan, et autres
Publié: (2024)
par: He, Ethan, et autres
Publié: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
par: Pan, Bowen, et autres
Publié: (2024)
par: Pan, Bowen, et autres
Publié: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
par: Li, Chong, et autres
Publié: (2025)
par: Li, Chong, et autres
Publié: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
SimpleStrat: Diversifying Language Model Generation with Stratification
par: Wong, Justin, et autres
Publié: (2024)
par: Wong, Justin, et autres
Publié: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
par: Liu, Zhili, et autres
Publié: (2024)
par: Liu, Zhili, et autres
Publié: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
par: Jiang, Fan, et autres
Publié: (2026)
par: Jiang, Fan, et autres
Publié: (2026)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
par: DeepSeek-AI, et autres
Publié: (2024)
par: DeepSeek-AI, et autres
Publié: (2024)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
par: Ho, Zheng Yi, et autres
Publié: (2024)
par: Ho, Zheng Yi, et autres
Publié: (2024)
LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling
par: Xie, Keqin
Publié: (2026)
par: Xie, Keqin
Publié: (2026)
Multi-Type Context-Aware Conversational Recommender Systems via Mixture-of-Experts
par: Zou, Jie, et autres
Publié: (2025)
par: Zou, Jie, et autres
Publié: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
par: Kim, Junhyuck, et autres
Publié: (2026)
par: Kim, Junhyuck, et autres
Publié: (2026)
OLMoE: Open Mixture-of-Experts Language Models
par: Muennighoff, Niklas, et autres
Publié: (2024)
par: Muennighoff, Niklas, et autres
Publié: (2024)
Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
par: Ternovtsii, Ivan, et autres
Publié: (2025)
par: Ternovtsii, Ivan, et autres
Publié: (2025)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
par: Lee, Rhui Dih, et autres
Publié: (2024)
par: Lee, Rhui Dih, et autres
Publié: (2024)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
par: Zhang, Yichi, et autres
Publié: (2024)
par: Zhang, Yichi, et autres
Publié: (2024)
Multilingual Routing in Mixture-of-Experts
par: Bandarkar, Lucas, et autres
Publié: (2025)
par: Bandarkar, Lucas, et autres
Publié: (2025)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
par: Luo, Haotian, et autres
Publié: (2025)
par: Luo, Haotian, et autres
Publié: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
par: Kim, Gyeongman, et autres
Publié: (2025)
par: Kim, Gyeongman, et autres
Publié: (2025)
Editing Knowledge Representation of Language Model via Rephrased Prefix Prompts
par: Cai, Yuchen, et autres
Publié: (2024)
par: Cai, Yuchen, et autres
Publié: (2024)
Aligning Large Language Models from Self-Reference AI Feedback with one General Principle
par: Bao, Rong, et autres
Publié: (2024)
par: Bao, Rong, et autres
Publié: (2024)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
par: Liu, Yilun, et autres
Publié: (2025)
par: Liu, Yilun, et autres
Publié: (2025)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
par: Li, Dengchun, et autres
Publié: (2024)
par: Li, Dengchun, et autres
Publié: (2024)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
par: Nakamura, Taishi, et autres
Publié: (2025)
par: Nakamura, Taishi, et autres
Publié: (2025)
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
par: GigaChat team, et autres
Publié: (2025)
par: GigaChat team, et autres
Publié: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning
par: Zhang, Yanfang, et autres
Publié: (2024)
par: Zhang, Yanfang, et autres
Publié: (2024)
Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning
par: Yue, Murong, et autres
Publié: (2023)
par: Yue, Murong, et autres
Publié: (2023)
SE-GCL: An Event-Based Simple and Effective Graph Contrastive Learning for Text Representation
par: Meng, Tao, et autres
Publié: (2024)
par: Meng, Tao, et autres
Publié: (2024)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
par: Zhang, Jianbin, et autres
Publié: (2025)
par: Zhang, Jianbin, et autres
Publié: (2025)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
par: Herbst, Jeremy, et autres
Publié: (2026)
par: Herbst, Jeremy, et autres
Publié: (2026)
Documents similaires
-
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
par: Xu, Ziyang, et autres
Publié: (2024) -
Revisiting Catastrophic Forgetting in Large Language Model Tuning
par: Li, Hongyu, et autres
Publié: (2024) -
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
par: Zan, Changtong, et autres
Publié: (2024) -
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
par: Lu, Qingyu, et autres
Publié: (2025) -
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
par: Wu, Yuchen, et autres
Publié: (2025)