Decoupled Relative Learning Rate Schedules
Fuente:
arXiv
Saved in:
| Main Authors: | Ludziejewski, Jan, Małaśnicki, Jan, Pióro, Maciej, Krutul, Michał, Ciebiera, Kamil, Stefaniak, Maciej, Krajewski, Jakub, Sankowski, Piotr, Cygan, Marek, Adamczewski, Kamil, Jaszczur, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023)
by: Antoniak, Szymon, et al.
Published: (2023)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023)
by: Nauman, Michal, et al.
Published: (2023)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
Shapley Pruning for Neural Network Compression
by: Adamczewski, Kamil, et al.
Published: (2024)
by: Adamczewski, Kamil, et al.
Published: (2024)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
RoboMorph: Evolving Robot Morphology using Large Language Models
by: Qiu, Kevin, et al.
Published: (2024)
by: Qiu, Kevin, et al.
Published: (2024)
Unifying Deep Stochastic Processes for Image Enhancement
by: Kozłowski, Wojciech, et al.
Published: (2026)
by: Kozłowski, Wojciech, et al.
Published: (2026)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
by: Kubaty, Piotr, et al.
Published: (2024)
by: Kubaty, Piotr, et al.
Published: (2024)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
by: Janusz, Mikołaj, et al.
Published: (2025)
by: Janusz, Mikołaj, et al.
Published: (2025)
CounterFlowNet: From Minimal Changes to Meaningful Counterfactual Explanations
by: Furman, Oleksii, et al.
Published: (2026)
by: Furman, Oleksii, et al.
Published: (2026)
SupResDiffGAN a new approach for the Super-Resolution task
by: Kopeć, Dawid, et al.
Published: (2025)
by: Kopeć, Dawid, et al.
Published: (2025)
EXALT: EXplainable ALgorithmic Tools for Optimization Problems
by: Bączek, Zuzanna, et al.
Published: (2025)
by: Bączek, Zuzanna, et al.
Published: (2025)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
by: Kuzina, Anna, et al.
Published: (2025)
by: Kuzina, Anna, et al.
Published: (2025)
FlySearch: Exploring how vision-language models explore
by: Pardyl, Adam, et al.
Published: (2025)
by: Pardyl, Adam, et al.
Published: (2025)
Unlearning-based sliding window for continual learning under concept drift
by: Wozniak, Michal, et al.
Published: (2026)
by: Wozniak, Michal, et al.
Published: (2026)
Differentially Private Neural Tangent Kernels for Privacy-Preserving Data Generation
by: Yang, Yilin, et al.
Published: (2023)
by: Yang, Yilin, et al.
Published: (2023)
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
by: Qiu, Kevin, et al.
Published: (2025)
by: Qiu, Kevin, et al.
Published: (2025)
Accessory pathway localization with probabilistic density maps generated by a mobile application: Assessment of a full pre‐excitation net‐vector method
by: Marek Jastrzębski, et al.
Published: (2024)
by: Marek Jastrzębski, et al.
Published: (2024)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
Joint Diffusion models in Continual Learning
by: Skierś, Paweł, et al.
Published: (2024)
by: Skierś, Paweł, et al.
Published: (2024)
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
by: Nauman, Michal, et al.
Published: (2025)
by: Nauman, Michal, et al.
Published: (2025)
Interplay between tie strength and neighbourhood topology in complex networks: Granovetter's theory and beyond
by: Mrowinski, Maciej J, et al.
Published: (2024)
by: Mrowinski, Maciej J, et al.
Published: (2024)
Anderson localization induced by structural disorder
by: Bhattacharjee, Sourav, et al.
Published: (2024)
by: Bhattacharjee, Sourav, et al.
Published: (2024)
Analysis of managerial behaviors in business management
by: Górka, Ernest, et al.
Published: (2025)
by: Górka, Ernest, et al.
Published: (2025)
Particle physics DL-simulation with control over generated data properties
by: Rogoziński, Karol, et al.
Published: (2024)
by: Rogoziński, Karol, et al.
Published: (2024)
The Minimax Rate of Second-Order Calibration
by: Ciosek, Kamil, et al.
Published: (2026)
by: Ciosek, Kamil, et al.
Published: (2026)
Hybrid spectral-spatial domain registration for nanometric tracking in digital in-line holographic microscopy
by: Kalinowski, Kamil, et al.
Published: (2026)
by: Kalinowski, Kamil, et al.
Published: (2026)
Estimation methods of Matrix-valued AR model
by: Kołodziejski, Kamil
Published: (2025)
by: Kołodziejski, Kamil
Published: (2025)
Three-Factor Learning in Spiking Neural Networks: An Overview of Methods and Trends from a Machine Learning Perspective
by: Mazurek, Szymon, et al.
Published: (2025)
by: Mazurek, Szymon, et al.
Published: (2025)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
Similar Items
-
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025) -
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025) -
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025) -
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024) -
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)