Bridging RL Theory and Practice with the Effective Horizon
Fuente:
arXiv
Saved in:
| Main Authors: | Laidlaw, Cassidy, Russell, Stuart, Dragan, Anca |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
AssistanceZero: Scalably Solving Assistance Games
by: Laidlaw, Cassidy, et al.
Published: (2025)
by: Laidlaw, Cassidy, et al.
Published: (2025)
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024)
by: Carroll, Micah, et al.
Published: (2024)
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
by: Feng, Dylan, et al.
Published: (2026)
by: Feng, Dylan, et al.
Published: (2026)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
by: Ye, Yaowen, et al.
Published: (2025)
by: Ye, Yaowen, et al.
Published: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
by: Siththaranjan, Anand, et al.
Published: (2023)
by: Siththaranjan, Anand, et al.
Published: (2023)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
by: Lang, Leon, et al.
Published: (2024)
by: Lang, Leon, et al.
Published: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
RL, but don't do anything I wouldn't do
by: Cohen, Michael K., et al.
Published: (2024)
by: Cohen, Michael K., et al.
Published: (2024)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Adversaries Can Misuse Combinations of Safe Models
by: Jones, Erik, et al.
Published: (2024)
by: Jones, Erik, et al.
Published: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
by: Wu, Ian, et al.
Published: (2026)
by: Wu, Ian, et al.
Published: (2026)
Horizon Reduction Makes RL Scalable
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
by: Myers, Vivek, et al.
Published: (2025)
by: Myers, Vivek, et al.
Published: (2025)
Bridging Theory and Practice in Crafting Robust Spiking Reservoirs
by: Freddi, Ruggero, et al.
Published: (2026)
by: Freddi, Ruggero, et al.
Published: (2026)
Aligning Robot and Human Representations
by: Bobu, Andreea, et al.
Published: (2023)
by: Bobu, Andreea, et al.
Published: (2023)
Training LLM Agents to Empower Humans
by: Ellis, Evan, et al.
Published: (2025)
by: Ellis, Evan, et al.
Published: (2025)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding
by: Xu, Jiamin, et al.
Published: (2026)
by: Xu, Jiamin, et al.
Published: (2026)
Feasibility-Aware Pessimistic Estimation: Toward Long-Horizon Safety in Offline RL
by: Tao, Zhikun
Published: (2025)
by: Tao, Zhikun
Published: (2025)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
by: Pan, Michelle, et al.
Published: (2024)
by: Pan, Michelle, et al.
Published: (2024)
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
by: Li, Yunfei, et al.
Published: (2025)
by: Li, Yunfei, et al.
Published: (2025)
Bridging Theory and Practice in Link Representation with Graph Neural Networks
by: Lachi, Veronica, et al.
Published: (2025)
by: Lachi, Veronica, et al.
Published: (2025)
Model-based RL as a Minimalist Approach to Horizon-Free and Second-Order Bounds
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
by: Zhu, Taojie, et al.
Published: (2026)
by: Zhu, Taojie, et al.
Published: (2026)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026)
by: Li, Yingru, et al.
Published: (2026)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
by: Su, Jianhai, et al.
Published: (2025)
by: Su, Jianhai, et al.
Published: (2025)
Offline Imitation Learning by Controlling the Effective Planning Horizon
by: Ahn, Hee-Jun, et al.
Published: (2024)
by: Ahn, Hee-Jun, et al.
Published: (2024)
On the Effective Horizon of Inverse Reinforcement Learning
by: Xu, Yiqing, et al.
Published: (2023)
by: Xu, Yiqing, et al.
Published: (2023)
Learning to Assist Humans without Inferring Rewards
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
Prior-Aligned Meta-RL: Thompson Sampling with Learned Priors and Guarantees in Finite-Horizon MDPs
by: Zhou, Runlin, et al.
Published: (2025)
by: Zhou, Runlin, et al.
Published: (2025)
On Representation Complexity of Model-based and Model-free Reinforcement Learning
by: Zhu, Hanlin, et al.
Published: (2023)
by: Zhu, Hanlin, et al.
Published: (2023)
Effective Interplay between Sparsity and Quantization: From Theory to Practice
by: Harma, Simla Burcu, et al.
Published: (2024)
by: Harma, Simla Burcu, et al.
Published: (2024)
Coalgebras for categorical deep learning: Representability and universal approximation
by: Mašulović, Dragan
Published: (2026)
by: Mašulović, Dragan
Published: (2026)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
by: Tomihari, Akiyoshi, et al.
Published: (2026)
by: Tomihari, Akiyoshi, et al.
Published: (2026)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
by: Choi, Jinwoo, et al.
Published: (2026)
by: Choi, Jinwoo, et al.
Published: (2026)
Randomness Helps Rigor: A Probabilistic Learning Rate Scheduler Bridging Theory and Deep Learning Practice
by: Devapriya, Dahlia, et al.
Published: (2024)
by: Devapriya, Dahlia, et al.
Published: (2024)
Similar Items
-
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023) -
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024) -
AssistanceZero: Scalably Solving Assistance Games
by: Laidlaw, Cassidy, et al.
Published: (2025) -
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024) -
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
by: Feng, Dylan, et al.
Published: (2026)