$μ$-Parametrization for Mixture of Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Małaśnicki, Jan, Ciebiera, Kamil, Boruń, Mateusz, Pióro, Maciej, Ludziejewski, Jan, Stefaniak, Maciej, Krutul, Michał, Jaszczur, Sebastian, Cygan, Marek, Adamczewski, Kamil, Krajewski, Jakub |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decoupled Relative Learning Rate Schedules
por: Ludziejewski, Jan, et al.
Publicado: (2025)
por: Ludziejewski, Jan, et al.
Publicado: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
por: Stefaniak, Maciej, et al.
Publicado: (2025)
por: Stefaniak, Maciej, et al.
Publicado: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
por: Ludziejewski, Jan, et al.
Publicado: (2025)
por: Ludziejewski, Jan, et al.
Publicado: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
por: Krajewski, Jakub, et al.
Publicado: (2024)
por: Krajewski, Jakub, et al.
Publicado: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
por: Antoniak, Szymon, et al.
Publicado: (2023)
por: Antoniak, Szymon, et al.
Publicado: (2023)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
por: Binkowski, Jakub, et al.
Publicado: (2026)
por: Binkowski, Jakub, et al.
Publicado: (2026)
A Case for Validation Buffer in Pessimistic Actor-Critic
por: Nauman, Michal, et al.
Publicado: (2024)
por: Nauman, Michal, et al.
Publicado: (2024)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
por: Nauman, Michal, et al.
Publicado: (2023)
por: Nauman, Michal, et al.
Publicado: (2023)
FlySearch: Exploring how vision-language models explore
por: Pardyl, Adam, et al.
Publicado: (2025)
por: Pardyl, Adam, et al.
Publicado: (2025)
RoboMorph: Evolving Robot Morphology using Large Language Models
por: Qiu, Kevin, et al.
Publicado: (2024)
por: Qiu, Kevin, et al.
Publicado: (2024)
Shapley Pruning for Neural Network Compression
por: Adamczewski, Kamil, et al.
Publicado: (2024)
por: Adamczewski, Kamil, et al.
Publicado: (2024)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
por: Nauman, Michal, et al.
Publicado: (2024)
por: Nauman, Michal, et al.
Publicado: (2024)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
por: Nauman, Michal, et al.
Publicado: (2024)
por: Nauman, Michal, et al.
Publicado: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
Reward-Conditioned Reinforcement Learning
por: Nauman, Michal, et al.
Publicado: (2026)
por: Nauman, Michal, et al.
Publicado: (2026)
MolPILE -- large-scale, diverse dataset for molecular representation learning
por: Adamczyk, Jakub, et al.
Publicado: (2025)
por: Adamczyk, Jakub, et al.
Publicado: (2025)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
por: Janusz, Mikołaj, et al.
Publicado: (2025)
por: Janusz, Mikołaj, et al.
Publicado: (2025)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
por: Kuzina, Anna, et al.
Publicado: (2025)
por: Kuzina, Anna, et al.
Publicado: (2025)
Unlearning-based sliding window for continual learning under concept drift
por: Wozniak, Michal, et al.
Publicado: (2026)
por: Wozniak, Michal, et al.
Publicado: (2026)
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
por: Qiu, Kevin, et al.
Publicado: (2025)
por: Qiu, Kevin, et al.
Publicado: (2025)
How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
por: Kubaty, Piotr, et al.
Publicado: (2024)
por: Kubaty, Piotr, et al.
Publicado: (2024)
Differentially Private Neural Tangent Kernels for Privacy-Preserving Data Generation
por: Yang, Yilin, et al.
Publicado: (2023)
por: Yang, Yilin, et al.
Publicado: (2023)
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
por: Nauman, Michal, et al.
Publicado: (2025)
por: Nauman, Michal, et al.
Publicado: (2025)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
por: Pardyl, Adam, et al.
Publicado: (2024)
por: Pardyl, Adam, et al.
Publicado: (2024)
TADA! Tuning Audio Diffusion Models through Activation Steering
por: Staniszewski, Łukasz, et al.
Publicado: (2026)
por: Staniszewski, Łukasz, et al.
Publicado: (2026)
Particle physics DL-simulation with control over generated data properties
por: Rogoziński, Karol, et al.
Publicado: (2024)
por: Rogoziński, Karol, et al.
Publicado: (2024)
Unifying Deep Stochastic Processes for Image Enhancement
por: Kozłowski, Wojciech, et al.
Publicado: (2026)
por: Kozłowski, Wojciech, et al.
Publicado: (2026)
CounterFlowNet: From Minimal Changes to Meaningful Counterfactual Explanations
por: Furman, Oleksii, et al.
Publicado: (2026)
por: Furman, Oleksii, et al.
Publicado: (2026)
Estimation methods of Matrix-valued AR model
por: Kołodziejski, Kamil
Publicado: (2025)
por: Kołodziejski, Kamil
Publicado: (2025)
Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts
por: Hendawy, Ahmed, et al.
Publicado: (2023)
por: Hendawy, Ahmed, et al.
Publicado: (2023)
Joint Diffusion models in Continual Learning
por: Skierś, Paweł, et al.
Publicado: (2024)
por: Skierś, Paweł, et al.
Publicado: (2024)
Membership and Dataset Inference Attacks on Large Audio Generative Models
por: Proboszcz, Jakub, et al.
Publicado: (2025)
por: Proboszcz, Jakub, et al.
Publicado: (2025)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
por: Wołczyk, Maciej, et al.
Publicado: (2024)
por: Wołczyk, Maciej, et al.
Publicado: (2024)
Graph Neural Networks on Graph Databases
por: Lopushanskyy, Dmytro, et al.
Publicado: (2024)
por: Lopushanskyy, Dmytro, et al.
Publicado: (2024)
Clustering with minimum spanning trees: How good can it be?
por: Gagolewski, Marek, et al.
Publicado: (2023)
por: Gagolewski, Marek, et al.
Publicado: (2023)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
por: Zawalski, Michał, et al.
Publicado: (2024)
por: Zawalski, Michał, et al.
Publicado: (2024)
Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
por: Chowdhury, Borun D
Publicado: (2026)
por: Chowdhury, Borun D
Publicado: (2026)
Vid2Sid: Videos Can Help Close the Sim2Real Gap
por: Qiu, Kevin, et al.
Publicado: (2026)
por: Qiu, Kevin, et al.
Publicado: (2026)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
por: Cywiński, Bartosz, et al.
Publicado: (2025)
por: Cywiński, Bartosz, et al.
Publicado: (2025)
Ejemplares similares
-
Decoupled Relative Learning Rate Schedules
por: Ludziejewski, Jan, et al.
Publicado: (2025) -
Projected Compression: Trainable Projection for Efficient Transformer Compression
por: Stefaniak, Maciej, et al.
Publicado: (2025) -
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
por: Ludziejewski, Jan, et al.
Publicado: (2025) -
Scaling Laws for Fine-Grained Mixture of Experts
por: Krajewski, Jakub, et al.
Publicado: (2024) -
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)