Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Zhiyuan, Guo, Qipeng, Fei, Zhaoye, Yin, Zhangyue, Zhou, Yunhua, Li, Linyang, Sun, Tianxiang, Yan, Hang, Lin, Dahua, Qiu, Xipeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Data-free Weight Compress and Denoise for Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2024)
di: Peng, Runyu, et al.
Pubblicazione: (2024)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
di: Wu, Haoze, et al.
Pubblicazione: (2024)
di: Wu, Haoze, et al.
Pubblicazione: (2024)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
di: Zhang, Yechen, et al.
Pubblicazione: (2026)
di: Zhang, Yechen, et al.
Pubblicazione: (2026)
How to Mitigate Overfitting in Weak-to-strong Generalization?
di: Shi, Junhao, et al.
Pubblicazione: (2025)
di: Shi, Junhao, et al.
Pubblicazione: (2025)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
di: Shi, Junhao, et al.
Pubblicazione: (2025)
di: Shi, Junhao, et al.
Pubblicazione: (2025)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
Unified Active Retrieval for Retrieval Augmented Generation
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
How to Set the Learning Rate for Large-Scale Pre-training?
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
Identifying Semantic Induction Heads to Understand In-Context Learning
di: Ren, Jie, et al.
Pubblicazione: (2024)
di: Ren, Jie, et al.
Pubblicazione: (2024)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
di: Lv, Kai, et al.
Pubblicazione: (2024)
di: Lv, Kai, et al.
Pubblicazione: (2024)
Dynamic and Generalizable Process Reward Modeling
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2024)
di: Yin, Zhangyue, et al.
Pubblicazione: (2024)
Code Needs Comments: Enhancing Code LLMs with Comment Augmentation
di: Song, Demin, et al.
Pubblicazione: (2024)
di: Song, Demin, et al.
Pubblicazione: (2024)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
di: Lv, Kai, et al.
Pubblicazione: (2023)
di: Lv, Kai, et al.
Pubblicazione: (2023)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
di: Wang, Yikun, et al.
Pubblicazione: (2025)
di: Wang, Yikun, et al.
Pubblicazione: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
di: Sun, Yu, et al.
Pubblicazione: (2024)
di: Sun, Yu, et al.
Pubblicazione: (2024)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
di: Lv, Kai, et al.
Pubblicazione: (2025)
di: Lv, Kai, et al.
Pubblicazione: (2025)
Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
di: Harvey, Daniel Fidel, et al.
Pubblicazione: (2025)
di: Harvey, Daniel Fidel, et al.
Pubblicazione: (2025)
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
di: Calanzone, Diego, et al.
Pubblicazione: (2025)
di: Calanzone, Diego, et al.
Pubblicazione: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
di: Ma, Wenhan, et al.
Pubblicazione: (2025)
di: Ma, Wenhan, et al.
Pubblicazione: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
di: Ye, Jiasheng, et al.
Pubblicazione: (2024)
di: Ye, Jiasheng, et al.
Pubblicazione: (2024)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
di: Wang, Xinghao, et al.
Pubblicazione: (2024)
di: Wang, Xinghao, et al.
Pubblicazione: (2024)
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
di: Li, Peiji, et al.
Pubblicazione: (2026)
di: Li, Peiji, et al.
Pubblicazione: (2026)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
FastMCTS: A Simple Sampling Strategy for Data Synthesis
di: Li, Peiji, et al.
Pubblicazione: (2025)
di: Li, Peiji, et al.
Pubblicazione: (2025)
Case2Code: Scalable Synthetic Data for Code Generation
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024) -
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024) -
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025) -
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2025) -
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023)