Scaling Laws for Fine-Grained Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Krajewski, Jakub, Ludziejewski, Jan, Adamczewski, Kamil, Pióro, Maciej, Krutul, Michał, Antoniak, Szymon, Ciebiera, Kamil, Król, Krystian, Odrzygóźdź, Tomasz, Sankowski, Piotr, Cygan, Marek, Jaszczur, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023)
by: Antoniak, Szymon, et al.
Published: (2023)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
Unifying Deep Stochastic Processes for Image Enhancement
by: Kozłowski, Wojciech, et al.
Published: (2026)
by: Kozłowski, Wojciech, et al.
Published: (2026)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
by: Janusz, Mikołaj, et al.
Published: (2025)
by: Janusz, Mikołaj, et al.
Published: (2025)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
by: Kubaty, Piotr, et al.
Published: (2024)
by: Kubaty, Piotr, et al.
Published: (2024)
ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
by: Będkowski, Patryk, et al.
Published: (2025)
by: Będkowski, Patryk, et al.
Published: (2025)
Shapley Pruning for Neural Network Compression
by: Adamczewski, Kamil, et al.
Published: (2024)
by: Adamczewski, Kamil, et al.
Published: (2024)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023)
by: Nauman, Michal, et al.
Published: (2023)
RoboMorph: Evolving Robot Morphology using Large Language Models
by: Qiu, Kevin, et al.
Published: (2024)
by: Qiu, Kevin, et al.
Published: (2024)
Robotics, AI and Criminal Law
by: Mamak, Kamil
Published: (2025)
by: Mamak, Kamil
Published: (2025)
Subgoal Search For Complex Reasoning Tasks
by: Czechowski, Konrad, et al.
Published: (2021)
by: Czechowski, Konrad, et al.
Published: (2021)
Conversations on Mind, Matter, and Mathematics. Jean-Pierre Changeux and Alain Connes. Edited and translated by M. B. DeBevoise. Princeton, New Jersey: Princeton University Press, 1995, 261 p.; 15 x 22 cm. Glossary + Index. Language: English. ISBN: 0-691-08759-8.
by: Marek Antoniak
Published: (2013)
by: Marek Antoniak
Published: (2013)
Solutions at vacuum and rarefaction waves in pressureless Euler alignment system
by: Cygan, Szymon, et al.
Published: (2024)
by: Cygan, Szymon, et al.
Published: (2024)
Perfect quantum protractors
by: Piotrak, Michał, et al.
Published: (2023)
by: Piotrak, Michał, et al.
Published: (2023)
Self-Supervised Event Representations: Towards Accurate, Real-Time Perception on SoC FPGAs
by: Jeziorek, Kamil, et al.
Published: (2025)
by: Jeziorek, Kamil, et al.
Published: (2025)
Imitating quantum probabilities: Beyond Bell's theorem and Tsirelson bounds
by: Czachor, Marek, et al.
Published: (2021)
by: Czachor, Marek, et al.
Published: (2021)
Hydrodynamical constraints on bubble wall velocity
by: Krajewski, Tomasz, et al.
Published: (2023)
by: Krajewski, Tomasz, et al.
Published: (2023)
Bubble-wall velocity in local thermal equilibrium: hydrodynamical simulations vs analytical treatment
by: Krajewski, Tomasz, et al.
Published: (2024)
by: Krajewski, Tomasz, et al.
Published: (2024)
Increasing the scalability of graph convolution for FPGA-implemented event-based vision
by: Wzorek, Piotr, et al.
Published: (2024)
by: Wzorek, Piotr, et al.
Published: (2024)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Heat kernels, intrinsic contractivity and ergodicity of discrete-time Markov chains killed by potentials
by: Cygan, Wojciech, et al.
Published: (2025)
by: Cygan, Wojciech, et al.
Published: (2025)
Fast and Precise: Adjusting Planning Horizon with Adaptive Subgoal Search
by: Zawalski, Michał, et al.
Published: (2022)
by: Zawalski, Michał, et al.
Published: (2022)
LLM generated responses to mitigate the impact of hate speech
by: Podolak, Jakub, et al.
Published: (2023)
by: Podolak, Jakub, et al.
Published: (2023)
Accessory pathway localization with probabilistic density maps generated by a mobile application: Assessment of a full pre‐excitation net‐vector method
by: Marek Jastrzębski, et al.
Published: (2024)
by: Marek Jastrzębski, et al.
Published: (2024)
Fully Dynamic Shortest Paths in Sparse Digraphs
by: Karczmarz, Adam, et al.
Published: (2024)
by: Karczmarz, Adam, et al.
Published: (2024)
Differentially Private Neural Tangent Kernels for Privacy-Preserving Data Generation
by: Yang, Yilin, et al.
Published: (2023)
by: Yang, Yilin, et al.
Published: (2023)
Online Multi-level Aggregation with Delays and Stochastic Arrivals
by: Mari, Mathieu, et al.
Published: (2024)
by: Mari, Mathieu, et al.
Published: (2024)
EXALT: EXplainable ALgorithmic Tools for Optimization Problems
by: Bączek, Zuzanna, et al.
Published: (2025)
by: Bączek, Zuzanna, et al.
Published: (2025)
Online matching with delays and stochastic arrival times
by: Mari, Mathieu, et al.
Published: (2022)
by: Mari, Mathieu, et al.
Published: (2022)
Cavity-enhanced spectroscopy in the deep cryogenic regime -- new hydrogen technologies for quantum sensing
by: Stankiewicz, Kamil, et al.
Published: (2025)
by: Stankiewicz, Kamil, et al.
Published: (2025)
Hybrid spectral-spatial domain registration for nanometric tracking in digital in-line holographic microscopy
by: Kalinowski, Kamil, et al.
Published: (2026)
by: Kalinowski, Kamil, et al.
Published: (2026)
Similar Items
-
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023) -
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024) -
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025) -
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025) -
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)