Mixtures of Experts Unlock Parameter Scaling for Deep RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Obando-Ceron, Johan, Sokar, Ghada, Willi, Timon, Lyle, Clare, Farebrother, Jesse, Foerster, Jakob, Dziugaite, Gintare Karolina, Precup, Doina, Castro, Pablo Samuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
von: Sepahvand, Nazanin Mohammadi, et al.
Veröffentlicht: (2026)
von: Sepahvand, Nazanin Mohammadi, et al.
Veröffentlicht: (2026)
Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Less is More: Undertraining Experts Improves Model Upcycling
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
CALE: Continuous Arcade Learning Environment
von: Farebrother, Jesse, et al.
Veröffentlicht: (2024)
von: Farebrother, Jesse, et al.
Veröffentlicht: (2024)
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
von: Sharma, Ekansh, et al.
Veröffentlicht: (2024)
von: Sharma, Ekansh, et al.
Veröffentlicht: (2024)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
von: Yang, Yu, et al.
Veröffentlicht: (2023)
von: Yang, Yu, et al.
Veröffentlicht: (2023)
Data Selection for Transfer Unlearning
von: Sepahvand, Nazanin Mohammadi, et al.
Veröffentlicht: (2024)
von: Sepahvand, Nazanin Mohammadi, et al.
Veröffentlicht: (2024)
Leveraging Function Space Aggregation for Federated Learning at Scale
von: Dhawan, Nikita, et al.
Veröffentlicht: (2023)
von: Dhawan, Nikita, et al.
Veröffentlicht: (2023)
Dataset Difficulty and the Role of Inductive Bias
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
von: Kwok, Devin, et al.
Veröffentlicht: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
von: Kleiman, Anat, et al.
Veröffentlicht: (2025)
von: Kleiman, Anat, et al.
Veröffentlicht: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
von: Baluta, Teodora, et al.
Veröffentlicht: (2024)
von: Baluta, Teodora, et al.
Veröffentlicht: (2024)
What Can Grokking Teach Us About Learning Under Nonstationarity?
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
In value-based deep reinforcement learning, a pruned network is a good network
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
The Danger Of Arrogance: Welfare Equilibra As A Solution To Stackelberg Self-Play In Non-Coincidental Games
von: Levi, Jake, et al.
Veröffentlicht: (2024)
von: Levi, Jake, et al.
Veröffentlicht: (2024)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
Analysing the Sample Complexity of Opponent Shaping
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
von: Jin, Tian, et al.
Veröffentlicht: (2025)
von: Jin, Tian, et al.
Veröffentlicht: (2025)
Diversity-Enriched Option-Critic
von: Kamat, Anand, et al.
Veröffentlicht: (2020)
von: Kamat, Anand, et al.
Veröffentlicht: (2020)
Functional Acceleration for Policy Mirror Descent
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
von: Alver, Safa, et al.
Veröffentlicht: (2022)
von: Alver, Safa, et al.
Veröffentlicht: (2022)
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
von: Attias, Idan, et al.
Veröffentlicht: (2024)
von: Attias, Idan, et al.
Veröffentlicht: (2024)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
von: Ohib, Riyasat, et al.
Veröffentlicht: (2024)
von: Ohib, Riyasat, et al.
Veröffentlicht: (2024)
Improved Localized Machine Unlearning Through the Lens of Memorization
von: Torkzadehmahani, Reihaneh, et al.
Veröffentlicht: (2024)
von: Torkzadehmahani, Reihaneh, et al.
Veröffentlicht: (2024)
The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
von: Mayor, Walter, et al.
Veröffentlicht: (2025)
von: Mayor, Walter, et al.
Veröffentlicht: (2025)
Simultaneous linear connectivity of neural networks modulo permutation
von: Sharma, Ekansh, et al.
Veröffentlicht: (2024)
von: Sharma, Ekansh, et al.
Veröffentlicht: (2024)
On Traceability in $\ell_p$ Stochastic Convex Optimization
von: Voitovych, Sasha, et al.
Veröffentlicht: (2025)
von: Voitovych, Sasha, et al.
Veröffentlicht: (2025)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
Balancing Plasticity and Stability with Fast and Slow Successor Features
von: Chua, Raymond, et al.
Veröffentlicht: (2026)
von: Chua, Raymond, et al.
Veröffentlicht: (2026)
On the Privacy of Selection Mechanisms with Gaussian Noise
von: Lebensold, Jonathan, et al.
Veröffentlicht: (2024)
von: Lebensold, Jonathan, et al.
Veröffentlicht: (2024)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2025)
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2025)
Scaling Opponent Shaping to High Dimensional Games
von: Khan, Akbir, et al.
Veröffentlicht: (2023)
von: Khan, Akbir, et al.
Veröffentlicht: (2023)
Torque-Aware Momentum
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024) -
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
von: Sokar, Ghada, et al.
Veröffentlicht: (2024) -
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025) -
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
von: Sepahvand, Nazanin Mohammadi, et al.
Veröffentlicht: (2026) -
Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)