Guardado en:
| Autores principales: | Abel, David, Barreto, André, Bowling, Michael, Dabney, Will, Dong, Shi, Hansen, Steven, Harutyunyan, Anna, Khetarpal, Khimya, Lyle, Clare, Pascanu, Razvan, Piliouras, Georgios, Precup, Doina, Richens, Jonathan, Rowland, Mark, Schaul, Tom, Singh, Satinder |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.04403 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Plasticity as the Mirror of Empowerment
por: Abel, David, et al.
Publicado: (2025)
por: Abel, David, et al.
Publicado: (2025)
Disentangling the Causes of Plasticity Loss in Neural Networks
por: Lyle, Clare, et al.
Publicado: (2024)
por: Lyle, Clare, et al.
Publicado: (2024)
Normalization and effective learning rates in reinforcement learning
por: Lyle, Clare, et al.
Publicado: (2024)
por: Lyle, Clare, et al.
Publicado: (2024)
Affordances Enable Partial World Modeling with LLMs
por: Khetarpal, Khimya, et al.
Publicado: (2026)
por: Khetarpal, Khimya, et al.
Publicado: (2026)
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
por: Beukman, Michael, et al.
Publicado: (2026)
por: Beukman, Michael, et al.
Publicado: (2026)
Optimizing Return Distributions with Distributional Dynamic Programming
por: Pires, Bernardo Ávila, et al.
Publicado: (2025)
por: Pires, Bernardo Ávila, et al.
Publicado: (2025)
Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
por: Cherif, Lynn, et al.
Publicado: (2025)
por: Cherif, Lynn, et al.
Publicado: (2025)
A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning
por: Khetarpal, Khimya, et al.
Publicado: (2024)
por: Khetarpal, Khimya, et al.
Publicado: (2024)
Capacity-Constrained Continual Learning
por: Wen, Zheng, et al.
Publicado: (2025)
por: Wen, Zheng, et al.
Publicado: (2025)
What Can Grokking Teach Us About Learning Under Nonstationarity?
por: Lyle, Clare, et al.
Publicado: (2025)
por: Lyle, Clare, et al.
Publicado: (2025)
Robust Intervention Learning from Emergency Stop Interventions
por: Pronovost, Ethan, et al.
Publicado: (2026)
por: Pronovost, Ethan, et al.
Publicado: (2026)
Boundless Socratic Learning with Language Games
por: Schaul, Tom
Publicado: (2024)
por: Schaul, Tom
Publicado: (2024)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
por: Rowland, Mark, et al.
Publicado: (2024)
por: Rowland, Mark, et al.
Publicado: (2024)
Diversity-Enriched Option-Critic
por: Kamat, Anand, et al.
Publicado: (2020)
por: Kamat, Anand, et al.
Publicado: (2020)
Functional Acceleration for Policy Mirror Descent
por: Chelu, Veronica, et al.
Publicado: (2024)
por: Chelu, Veronica, et al.
Publicado: (2024)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
por: Alver, Safa, et al.
Publicado: (2022)
por: Alver, Safa, et al.
Publicado: (2022)
Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning
por: Lawson, Daniel, et al.
Publicado: (2025)
por: Lawson, Daniel, et al.
Publicado: (2025)
Capturing Individual Human Preferences with Reward Features
por: Barreto, André, et al.
Publicado: (2025)
por: Barreto, André, et al.
Publicado: (2025)
General agents contain world models
por: Richens, Jonathan, et al.
Publicado: (2025)
por: Richens, Jonathan, et al.
Publicado: (2025)
Fine-Tuned In-Context Learners for Efficient Adaptation
por: Bornschein, Jorg, et al.
Publicado: (2025)
por: Bornschein, Jorg, et al.
Publicado: (2025)
Robust agents learn causal world models
por: Richens, Jonathan, et al.
Publicado: (2024)
por: Richens, Jonathan, et al.
Publicado: (2024)
Representation Learning via Non-Contrastive Mutual Information
por: Guo, Zhaohan Daniel, et al.
Publicado: (2025)
por: Guo, Zhaohan Daniel, et al.
Publicado: (2025)
Balancing Plasticity and Stability with Fast and Slow Successor Features
por: Chua, Raymond, et al.
Publicado: (2026)
por: Chua, Raymond, et al.
Publicado: (2026)
On the Privacy of Selection Mechanisms with Gaussian Noise
por: Lebensold, Jonathan, et al.
Publicado: (2024)
por: Lebensold, Jonathan, et al.
Publicado: (2024)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
por: Galashov, Alexandre, et al.
Publicado: (2024)
por: Galashov, Alexandre, et al.
Publicado: (2024)
A Distributional Analogue to the Successor Representation
por: Wiltzer, Harley, et al.
Publicado: (2024)
por: Wiltzer, Harley, et al.
Publicado: (2024)
Conditions on Preference Relations that Guarantee the Existence of Optimal Policies
por: Carr, Jonathan Colaço, et al.
Publicado: (2023)
por: Carr, Jonathan Colaço, et al.
Publicado: (2023)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
por: Arnob, Samin Yeasar, et al.
Publicado: (2025)
por: Arnob, Samin Yeasar, et al.
Publicado: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
por: Jain, Arushi, et al.
Publicado: (2024)
por: Jain, Arushi, et al.
Publicado: (2024)
Fluid-Agent Reinforcement Learning
por: Sharma, Shishir, et al.
Publicado: (2026)
por: Sharma, Shishir, et al.
Publicado: (2026)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
por: Alver, Safa, et al.
Publicado: (2024)
por: Alver, Safa, et al.
Publicado: (2024)
Lattice: Learning to Efficiently Compress the Memory
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
por: Fan, Simin, et al.
Publicado: (2024)
por: Fan, Simin, et al.
Publicado: (2024)
Toward Human-AI Alignment in Large-Scale Multi-Player Games
por: Sharma, Sugandha, et al.
Publicado: (2024)
por: Sharma, Sugandha, et al.
Publicado: (2024)
The Limits of Predicting Agents from Behaviour
por: Bellot, Alexis, et al.
Publicado: (2025)
por: Bellot, Alexis, et al.
Publicado: (2025)
An Analysis of Quantile Temporal-Difference Learning
por: Rowland, Mark, et al.
Publicado: (2023)
por: Rowland, Mark, et al.
Publicado: (2023)
Parseval Regularization for Continual Reinforcement Learning
por: Chung, Wesley, et al.
Publicado: (2024)
por: Chung, Wesley, et al.
Publicado: (2024)
Relative Trajectory Balance is equivalent to Trust-PCL
por: Deleu, Tristan, et al.
Publicado: (2025)
por: Deleu, Tristan, et al.
Publicado: (2025)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
por: Li, Qinyu, et al.
Publicado: (2025)
por: Li, Qinyu, et al.
Publicado: (2025)
Meta-learning how to Share Credit among Macro-Actions
por: Hosu, Ionel-Alexandru, et al.
Publicado: (2025)
por: Hosu, Ionel-Alexandru, et al.
Publicado: (2025)
Ejemplares similares
-
Plasticity as the Mirror of Empowerment
por: Abel, David, et al.
Publicado: (2025) -
Disentangling the Causes of Plasticity Loss in Neural Networks
por: Lyle, Clare, et al.
Publicado: (2024) -
Normalization and effective learning rates in reinforcement learning
por: Lyle, Clare, et al.
Publicado: (2024) -
Affordances Enable Partial World Modeling with LLMs
por: Khetarpal, Khimya, et al.
Publicado: (2026) -
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
por: Beukman, Michael, et al.
Publicado: (2026)