Theory on Mixture-of-Experts in Continual Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Hongbo, Lin, Sen, Duan, Lingjie, Liang, Yingbin, Shroff, Ness B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
di: Shi, Ming, et al.
Pubblicazione: (2023)
di: Shi, Ming, et al.
Pubblicazione: (2023)
Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
di: Li, Hongbo, et al.
Pubblicazione: (2025)
di: Li, Hongbo, et al.
Pubblicazione: (2025)
Learnable Chernoff Baselines for Inference-Time Alignment
di: Madhow, Sunil, et al.
Pubblicazione: (2026)
di: Madhow, Sunil, et al.
Pubblicazione: (2026)
Theory of Mixture-of-Experts for Mobile Edge Computing
di: Li, Hongbo, et al.
Pubblicazione: (2024)
di: Li, Hongbo, et al.
Pubblicazione: (2024)
Provable In-Context Learning of Nonlinear Regression with Transformers
di: Li, Hongbo, et al.
Pubblicazione: (2025)
di: Li, Hongbo, et al.
Pubblicazione: (2025)
A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications
di: Mu, Siyuan, et al.
Pubblicazione: (2025)
di: Mu, Siyuan, et al.
Pubblicazione: (2025)
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
di: Deng, Junze, et al.
Pubblicazione: (2025)
di: Deng, Junze, et al.
Pubblicazione: (2025)
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
di: Li, Yining, et al.
Pubblicazione: (2026)
di: Li, Yining, et al.
Pubblicazione: (2026)
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
di: Hao, Shugang, et al.
Pubblicazione: (2025)
di: Hao, Shugang, et al.
Pubblicazione: (2025)
Monitoring State Transitions in Markovian Systems with Sampling Cost
di: Saurav, Kumar, et al.
Pubblicazione: (2025)
di: Saurav, Kumar, et al.
Pubblicazione: (2025)
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
di: Shi, Ming, et al.
Pubblicazione: (2026)
di: Shi, Ming, et al.
Pubblicazione: (2026)
Constraint-Rectified Training for Efficient Chain-of-Thought
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
Sharp Convergence Rates for Masked Diffusion Models
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
Broadening Target Distributions for Accelerated Diffusion Models via a Novel Analysis Approach
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
di: Liang, Yuchen, et al.
Pubblicazione: (2025)
di: Liang, Yuchen, et al.
Pubblicazione: (2025)
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2025)
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2025)
Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?
di: Ju, Peizhong, et al.
Pubblicazione: (2024)
di: Ju, Peizhong, et al.
Pubblicazione: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
di: He, Yifei, et al.
Pubblicazione: (2025)
di: He, Yifei, et al.
Pubblicazione: (2025)
Mixtures of SubExperts for Large Language Continual Learning
di: Kang, Haeyong
Pubblicazione: (2025)
di: Kang, Haeyong
Pubblicazione: (2025)
Mixture-of-Experts Meets In-Context Reinforcement Learning
di: Wu, Wenhao, et al.
Pubblicazione: (2025)
di: Wu, Wenhao, et al.
Pubblicazione: (2025)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
di: Li, Gang, et al.
Pubblicazione: (2024)
di: Li, Gang, et al.
Pubblicazione: (2024)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
di: Xie, Zhen-Hao, et al.
Pubblicazione: (2026)
di: Xie, Zhen-Hao, et al.
Pubblicazione: (2026)
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
di: Liang, Jingcong, et al.
Pubblicazione: (2025)
di: Liang, Jingcong, et al.
Pubblicazione: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
di: Chen, Keru, et al.
Pubblicazione: (2026)
di: Chen, Keru, et al.
Pubblicazione: (2026)
Aggregation Alignment for Federated Learning with Mixture-of-Experts under Data Heterogeneity
di: Fang, Zihan, et al.
Pubblicazione: (2026)
di: Fang, Zihan, et al.
Pubblicazione: (2026)
Mixture of Raytraced Experts
di: Perin, Andrea, et al.
Pubblicazione: (2025)
di: Perin, Andrea, et al.
Pubblicazione: (2025)
Mixture of Experts in a Mixture of RL settings
di: Willi, Timon, et al.
Pubblicazione: (2024)
di: Willi, Timon, et al.
Pubblicazione: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026)
di: Madan, Vivan, et al.
Pubblicazione: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
di: Jin, Peng, et al.
Pubblicazione: (2024)
di: Jin, Peng, et al.
Pubblicazione: (2024)
In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
di: Lin, Shuning, et al.
Pubblicazione: (2025)
di: Lin, Shuning, et al.
Pubblicazione: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Mixture of Diverse Size Experts
di: Sun, Manxi, et al.
Pubblicazione: (2024)
di: Sun, Manxi, et al.
Pubblicazione: (2024)
Mixture of A Million Experts
di: He, Xu Owen
Pubblicazione: (2024)
di: He, Xu Owen
Pubblicazione: (2024)
Sparsity and Superposition in Mixture of Experts
di: Chaudhari, Marmik, et al.
Pubblicazione: (2025)
di: Chaudhari, Marmik, et al.
Pubblicazione: (2025)
Mixture of Concept Bottleneck Experts
di: De Santis, Francesco, et al.
Pubblicazione: (2026)
di: De Santis, Francesco, et al.
Pubblicazione: (2026)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
di: Yang, Cheng, et al.
Pubblicazione: (2024)
di: Yang, Cheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
di: Shi, Ming, et al.
Pubblicazione: (2023) -
Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
di: Li, Hongbo, et al.
Pubblicazione: (2025) -
Learnable Chernoff Baselines for Inference-Time Alignment
di: Madhow, Sunil, et al.
Pubblicazione: (2026) -
Theory of Mixture-of-Experts for Mobile Edge Computing
di: Li, Hongbo, et al.
Pubblicazione: (2024) -
Provable In-Context Learning of Nonlinear Regression with Transformers
di: Li, Hongbo, et al.
Pubblicazione: (2025)