Code as Reward: Empowering Reinforcement Learning with VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Venuto, David, Islam, Sami Nur, Klissarov, Martin, Precup, Doina, Yang, Sherry, Anand, Ankit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
by: Ishfaq, Haque, et al.
Published: (2025)
by: Ishfaq, Haque, et al.
Published: (2025)
Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
by: Cherif, Lynn, et al.
Published: (2025)
by: Cherif, Lynn, et al.
Published: (2025)
Diversity-Enriched Option-Critic
by: Kamat, Anand, et al.
Published: (2020)
by: Kamat, Anand, et al.
Published: (2020)
Parseval Regularization for Continual Reinforcement Learning
by: Chung, Wesley, et al.
Published: (2024)
by: Chung, Wesley, et al.
Published: (2024)
Fluid-Agent Reinforcement Learning
by: Sharma, Shishir, et al.
Published: (2026)
by: Sharma, Shishir, et al.
Published: (2026)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
by: Alver, Safa, et al.
Published: (2024)
by: Alver, Safa, et al.
Published: (2024)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Functional Acceleration for Policy Mirror Descent
by: Chelu, Veronica, et al.
Published: (2024)
by: Chelu, Veronica, et al.
Published: (2024)
A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
by: Alver, Safa, et al.
Published: (2022)
by: Alver, Safa, et al.
Published: (2022)
Fairness in Reinforcement Learning with Bisimulation Metrics
by: Rezaei-Shoshtari, Sahand, et al.
Published: (2024)
by: Rezaei-Shoshtari, Sahand, et al.
Published: (2024)
Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
by: Zhang, Shuyuan, et al.
Published: (2025)
by: Zhang, Shuyuan, et al.
Published: (2025)
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
by: Carr, Jonathan Colaço, et al.
Published: (2026)
by: Carr, Jonathan Colaço, et al.
Published: (2026)
Balancing Plasticity and Stability with Fast and Slow Successor Features
by: Chua, Raymond, et al.
Published: (2026)
by: Chua, Raymond, et al.
Published: (2026)
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
by: Luo, Ziyan, et al.
Published: (2025)
by: Luo, Ziyan, et al.
Published: (2025)
Offline Multitask Representation Learning for Reinforcement Learning
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
On the Privacy of Selection Mechanisms with Gaussian Noise
by: Lebensold, Jonathan, et al.
Published: (2024)
by: Lebensold, Jonathan, et al.
Published: (2024)
Conditions on Preference Relations that Guarantee the Existence of Optimal Policies
by: Carr, Jonathan Colaço, et al.
Published: (2023)
by: Carr, Jonathan Colaço, et al.
Published: (2023)
Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement Learning
by: Zhao, Mingde, et al.
Published: (2023)
by: Zhao, Mingde, et al.
Published: (2023)
Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning
by: Klissarov, Martin, et al.
Published: (2025)
by: Klissarov, Martin, et al.
Published: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
Relative Trajectory Balance is equivalent to Trust-PCL
by: Deleu, Tristan, et al.
Published: (2025)
by: Deleu, Tristan, et al.
Published: (2025)
Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo
by: Ishfaq, Haque, et al.
Published: (2023)
by: Ishfaq, Haque, et al.
Published: (2023)
Learning Successor Features the Simple Way
by: Chua, Raymond, et al.
Published: (2024)
by: Chua, Raymond, et al.
Published: (2024)
Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards
by: Jolfaei, Erfan Aghadavoodi, et al.
Published: (2026)
by: Jolfaei, Erfan Aghadavoodi, et al.
Published: (2026)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
by: Klissarov, Martin, et al.
Published: (2024)
by: Klissarov, Martin, et al.
Published: (2024)
Capacity-Constrained Continual Learning
by: Wen, Zheng, et al.
Published: (2025)
by: Wen, Zheng, et al.
Published: (2025)
Discrete Probabilistic Inference as Control in Multi-path Environments
by: Deleu, Tristan, et al.
Published: (2024)
by: Deleu, Tristan, et al.
Published: (2024)
Capturing Individual Human Preferences with Reward Features
by: Barreto, André, et al.
Published: (2025)
by: Barreto, André, et al.
Published: (2025)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
by: Panangaden, Prakash, et al.
Published: (2023)
by: Panangaden, Prakash, et al.
Published: (2023)
QGFN: Controllable Greediness with Action Values
by: Lau, Elaine, et al.
Published: (2024)
by: Lau, Elaine, et al.
Published: (2024)
Robust Reward Modeling via Causal Rubrics
by: Srivastava, Pragya, et al.
Published: (2025)
by: Srivastava, Pragya, et al.
Published: (2025)
Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation
by: Patil, Gandharv, et al.
Published: (2022)
by: Patil, Gandharv, et al.
Published: (2022)
Plasticity as the Mirror of Empowerment
by: Abel, David, et al.
Published: (2025)
by: Abel, David, et al.
Published: (2025)
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
by: McCracken, Gavin, et al.
Published: (2025)
by: McCracken, Gavin, et al.
Published: (2025)
Reinforcement Learning for Machine Learning Engineering Agents
by: Yang, Sherry, et al.
Published: (2025)
by: Yang, Sherry, et al.
Published: (2025)
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
by: Zhu, Speed, et al.
Published: (2025)
by: Zhu, Speed, et al.
Published: (2025)
Effective Protein-Protein Interaction Exploration with PPIretrieval
by: Hua, Chenqing, et al.
Published: (2024)
by: Hua, Chenqing, et al.
Published: (2024)
Mitigating Downstream Model Risks via Model Provenance
by: Wang, Keyu, et al.
Published: (2024)
by: Wang, Keyu, et al.
Published: (2024)
Similar Items
-
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning
by: Ishfaq, Haque, et al.
Published: (2025) -
Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
by: Cherif, Lynn, et al.
Published: (2025) -
Diversity-Enriched Option-Critic
by: Kamat, Anand, et al.
Published: (2020) -
Parseval Regularization for Continual Reinforcement Learning
by: Chung, Wesley, et al.
Published: (2024) -
Fluid-Agent Reinforcement Learning
by: Sharma, Shishir, et al.
Published: (2026)