Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
Fuente:
arXiv
Saved in:
| Main Authors: | Nauman, Michal, Ostaszewski, Mateusz, Jankowski, Krzysztof, Miłoś, Piotr, Cygan, Marek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
by: Nauman, Michal, et al.
Published: (2025)
by: Nauman, Michal, et al.
Published: (2025)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023)
by: Nauman, Michal, et al.
Published: (2023)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
Learning Continually by Spectral Regularization
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
When Does Non-Uniform Replay Matter in Reinforcement Learning?
by: Korniak, Michal, et al.
Published: (2026)
by: Korniak, Michal, et al.
Published: (2026)
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
by: Bortkiewicz, Michał, et al.
Published: (2025)
by: Bortkiewicz, Michał, et al.
Published: (2025)
Amortized Causal Discovery with Prior-Fitted Networks
by: Sypniewski, Mateusz, et al.
Published: (2025)
by: Sypniewski, Mateusz, et al.
Published: (2025)
Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
by: Surdej, Rafał, et al.
Published: (2025)
by: Surdej, Rafał, et al.
Published: (2025)
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
by: Qiu, Kevin, et al.
Published: (2025)
by: Qiu, Kevin, et al.
Published: (2025)
A Python library for efficient computation of molecular fingerprints
by: Szafarczyk, Michał, et al.
Published: (2024)
by: Szafarczyk, Michał, et al.
Published: (2024)
Optimistic Policy Regularization
by: Pham, Mai, et al.
Published: (2026)
by: Pham, Mai, et al.
Published: (2026)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
by: Chen, Yimeng, et al.
Published: (2025)
by: Chen, Yimeng, et al.
Published: (2025)
Since Faithfulness Fails: The Performance Limits of Neural Causal Discovery
by: Olko, Mateusz, et al.
Published: (2025)
by: Olko, Mateusz, et al.
Published: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
RoboMorph: Evolving Robot Morphology using Large Language Models
by: Qiu, Kevin, et al.
Published: (2024)
by: Qiu, Kevin, et al.
Published: (2024)
Solving cold start in news recommendations: a RippleNet-based system for large scale media outlet
by: Radziszewski, Karol, et al.
Published: (2025)
by: Radziszewski, Karol, et al.
Published: (2025)
Subgoal Search For Complex Reasoning Tasks
by: Czechowski, Konrad, et al.
Published: (2021)
by: Czechowski, Konrad, et al.
Published: (2021)
RetroGFN: Diverse and Feasible Retrosynthesis using GFlowNets
by: Gaiński, Piotr, et al.
Published: (2024)
by: Gaiński, Piotr, et al.
Published: (2024)
FlySearch: Exploring how vision-language models explore
by: Pardyl, Adam, et al.
Published: (2025)
by: Pardyl, Adam, et al.
Published: (2025)
Vid2Sid: Videos Can Help Close the Sim2Real Gap
by: Qiu, Kevin, et al.
Published: (2026)
by: Qiu, Kevin, et al.
Published: (2026)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Scikit-fingerprints: easy and efficient computation of molecular fingerprints in Python
by: Adamczyk, Jakub, et al.
Published: (2024)
by: Adamczyk, Jakub, et al.
Published: (2024)
Contrastive Representations for Temporal Reasoning
by: Ziarko, Alicja, et al.
Published: (2025)
by: Ziarko, Alicja, et al.
Published: (2025)
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Unlearning-based sliding window for continual learning under concept drift
by: Wozniak, Michal, et al.
Published: (2026)
by: Wozniak, Michal, et al.
Published: (2026)
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025)
by: Masarczyk, Wojciech, et al.
Published: (2025)
Identifying Super Spreaders in Multilayer Networks
by: Czuba, Michał, et al.
Published: (2025)
by: Czuba, Michał, et al.
Published: (2025)
What Does Flow Matching Bring To TD Learning?
by: Agrawalla, Bhavya, et al.
Published: (2026)
by: Agrawalla, Bhavya, et al.
Published: (2026)
Feature importance analysis for patient management decisions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Distance metric learning for conditional anomaly detection
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Trust Your $\nabla$: Gradient-based Intervention Targeting for Causal Discovery
by: Olko, Mateusz, et al.
Published: (2022)
by: Olko, Mateusz, et al.
Published: (2022)
tsGT: Stochastic Time Series Modeling With Transformer
by: Kuciński, Łukasz, et al.
Published: (2024)
by: Kuciński, Łukasz, et al.
Published: (2024)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
Clustering-based hard negative sampling for supervised contrastive speaker verification
by: Masztalski, Piotr, et al.
Published: (2025)
by: Masztalski, Piotr, et al.
Published: (2025)
Similar Items
-
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024) -
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024) -
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
by: Nauman, Michal, et al.
Published: (2025) -
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023) -
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)