Guardado en:
| Autores principales: | Barreto, André, Dumoulin, Vincent, Mao, Yiran, Rowland, Mark, Perez-Nieves, Nicolas, Shahriari, Bobak, Dauphin, Yann, Precup, Doina, Larochelle, Hugo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2503.17338 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A density estimation perspective on learning from pairwise human preferences
por: Dumoulin, Vincent, et al.
Publicado: (2023)
por: Dumoulin, Vincent, et al.
Publicado: (2023)
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
por: Sepahvand, Nazanin Mohammadi, et al.
Publicado: (2026)
por: Sepahvand, Nazanin Mohammadi, et al.
Publicado: (2026)
Conditions on Preference Relations that Guarantee the Existence of Optimal Policies
por: Carr, Jonathan Colaço, et al.
Publicado: (2023)
por: Carr, Jonathan Colaço, et al.
Publicado: (2023)
Balancing Plasticity and Stability with Fast and Slow Successor Features
por: Chua, Raymond, et al.
Publicado: (2026)
por: Chua, Raymond, et al.
Publicado: (2026)
Diversity-Enriched Option-Critic
por: Kamat, Anand, et al.
Publicado: (2020)
por: Kamat, Anand, et al.
Publicado: (2020)
Functional Acceleration for Policy Mirror Descent
por: Chelu, Veronica, et al.
Publicado: (2024)
por: Chelu, Veronica, et al.
Publicado: (2024)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
por: Alver, Safa, et al.
Publicado: (2022)
por: Alver, Safa, et al.
Publicado: (2022)
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
por: Carr, Jonathan Colaço, et al.
Publicado: (2026)
por: Carr, Jonathan Colaço, et al.
Publicado: (2026)
On the Privacy of Selection Mechanisms with Gaussian Noise
por: Lebensold, Jonathan, et al.
Publicado: (2024)
por: Lebensold, Jonathan, et al.
Publicado: (2024)
Code as Reward: Empowering Reinforcement Learning with VLMs
por: Venuto, David, et al.
Publicado: (2024)
por: Venuto, David, et al.
Publicado: (2024)
Learning Successor Features the Simple Way
por: Chua, Raymond, et al.
Publicado: (2024)
por: Chua, Raymond, et al.
Publicado: (2024)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
por: Arnob, Samin Yeasar, et al.
Publicado: (2025)
por: Arnob, Samin Yeasar, et al.
Publicado: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
por: Jain, Arushi, et al.
Publicado: (2024)
por: Jain, Arushi, et al.
Publicado: (2024)
Fluid-Agent Reinforcement Learning
por: Sharma, Shishir, et al.
Publicado: (2026)
por: Sharma, Shishir, et al.
Publicado: (2026)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
por: Alver, Safa, et al.
Publicado: (2024)
por: Alver, Safa, et al.
Publicado: (2024)
Parseval Regularization for Continual Reinforcement Learning
por: Chung, Wesley, et al.
Publicado: (2024)
por: Chung, Wesley, et al.
Publicado: (2024)
Relative Trajectory Balance is equivalent to Trust-PCL
por: Deleu, Tristan, et al.
Publicado: (2025)
por: Deleu, Tristan, et al.
Publicado: (2025)
Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
por: McCracken, Gavin, et al.
Publicado: (2025)
por: McCracken, Gavin, et al.
Publicado: (2025)
Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
por: Zhang, Shuyuan, et al.
Publicado: (2025)
por: Zhang, Shuyuan, et al.
Publicado: (2025)
SCAR: Shapley Credit Assignment for More Efficient RLHF
por: Cao, Meng, et al.
Publicado: (2025)
por: Cao, Meng, et al.
Publicado: (2025)
Capacity-Constrained Continual Learning
por: Wen, Zheng, et al.
Publicado: (2025)
por: Wen, Zheng, et al.
Publicado: (2025)
Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation
por: Patil, Gandharv, et al.
Publicado: (2022)
por: Patil, Gandharv, et al.
Publicado: (2022)
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
por: Ishfaq, Haque, et al.
Publicado: (2025)
por: Ishfaq, Haque, et al.
Publicado: (2025)
Plasticity as the Mirror of Empowerment
por: Abel, David, et al.
Publicado: (2025)
por: Abel, David, et al.
Publicado: (2025)
Agency Is Frame-Dependent
por: Abel, David, et al.
Publicado: (2025)
por: Abel, David, et al.
Publicado: (2025)
Discrete Probabilistic Inference as Control in Multi-path Environments
por: Deleu, Tristan, et al.
Publicado: (2024)
por: Deleu, Tristan, et al.
Publicado: (2024)
Rejecting Hallucinated State Targets during Planning
por: Zhao, Mingde, et al.
Publicado: (2024)
por: Zhao, Mingde, et al.
Publicado: (2024)
Effective Protein-Protein Interaction Exploration with PPIretrieval
por: Hua, Chenqing, et al.
Publicado: (2024)
por: Hua, Chenqing, et al.
Publicado: (2024)
Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input
por: Peng, Andi, et al.
Publicado: (2024)
por: Peng, Andi, et al.
Publicado: (2024)
Zolus kauriensis Larochelle & Larivière & Larochelle & Larivière 2017, new species
por: Larochelle, et al.
Publicado: (2017)
por: Larochelle, et al.
Publicado: (2017)
Maungazolus ranatungae Larochelle & Larivière & Larochelle & Larivière 2017, new species
por: Larochelle, et al.
Publicado: (2017)
por: Larochelle, et al.
Publicado: (2017)
CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
por: Budaghyan, David, et al.
Publicado: (2023)
por: Budaghyan, David, et al.
Publicado: (2023)
Fairness in Reinforcement Learning with Bisimulation Metrics
por: Rezaei-Shoshtari, Sahand, et al.
Publicado: (2024)
por: Rezaei-Shoshtari, Sahand, et al.
Publicado: (2024)
QGFN: Controllable Greediness with Action Values
por: Lau, Elaine, et al.
Publicado: (2024)
por: Lau, Elaine, et al.
Publicado: (2024)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
por: Panangaden, Prakash, et al.
Publicado: (2023)
por: Panangaden, Prakash, et al.
Publicado: (2023)
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
por: Luo, Ziyan, et al.
Publicado: (2025)
por: Luo, Ziyan, et al.
Publicado: (2025)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
por: Padmakumar, Vishakh, et al.
Publicado: (2024)
por: Padmakumar, Vishakh, et al.
Publicado: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
por: Liu, Fengyuan, et al.
Publicado: (2026)
por: Liu, Fengyuan, et al.
Publicado: (2026)
Robust Reward Modeling via Causal Rubrics
por: Srivastava, Pragya, et al.
Publicado: (2025)
por: Srivastava, Pragya, et al.
Publicado: (2025)
On the Limits of Multi-modal Meta-Learning with Auxiliary Task Modulation Using Conditional Batch Normalization
por: Armengol-Estapé, Jordi, et al.
Publicado: (2024)
por: Armengol-Estapé, Jordi, et al.
Publicado: (2024)
Ejemplares similares
-
A density estimation perspective on learning from pairwise human preferences
por: Dumoulin, Vincent, et al.
Publicado: (2023) -
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
por: Sepahvand, Nazanin Mohammadi, et al.
Publicado: (2026) -
Conditions on Preference Relations that Guarantee the Existence of Optimal Policies
por: Carr, Jonathan Colaço, et al.
Publicado: (2023) -
Balancing Plasticity and Stability with Fast and Slow Successor Features
por: Chua, Raymond, et al.
Publicado: (2026) -
Diversity-Enriched Option-Critic
por: Kamat, Anand, et al.
Publicado: (2020)