Saved in:
| Main Authors: | Hussing, Marcel, Voelcker, Claas, Gilitschenski, Igor, Farahmand, Amir-massoud, Eaton, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.05996 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
by: Voelcker, Claas A, et al.
Published: (2023)
by: Voelcker, Claas A, et al.
Published: (2023)
When does Self-Prediction help? Understanding Auxiliary Tasks in Reinforcement Learning
by: Voelcker, Claas, et al.
Published: (2024)
by: Voelcker, Claas, et al.
Published: (2024)
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Behavior-Consistent Deep Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2026)
by: Hussing, Marcel, et al.
Published: (2026)
Calibrated Value-Aware Model Learning with Probabilistic Environment Models
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Can we hop in general? A discussion of benchmark selection and design using the Hopper environment
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024)
by: Bedaywi, Mark, et al.
Published: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
Robotic Manipulation Datasets for Offline Compositional Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2023)
by: Hussing, Marcel, et al.
Published: (2023)
Model Agreement via Anchoring
by: Eaton, Eric, et al.
Published: (2026)
by: Eaton, Eric, et al.
Published: (2026)
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
by: Opryshko, Evgenii, et al.
Published: (2025)
by: Opryshko, Evgenii, et al.
Published: (2025)
Oracle-Efficient Reinforcement Learning for Max Value Ensembles
by: Hussing, Marcel, et al.
Published: (2024)
by: Hussing, Marcel, et al.
Published: (2024)
Distributed Continual Learning
by: Le, Long, et al.
Published: (2024)
by: Le, Long, et al.
Published: (2024)
Deflated Dynamics Value Iteration
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025)
by: Ma, Avery, et al.
Published: (2025)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
by: Zheng, Shuhong, et al.
Published: (2025)
by: Zheng, Shuhong, et al.
Published: (2025)
Efficient and Accurate Optimal Transport with Mirror Descent and Conjugate Gradients
by: Kemertas, Mete, et al.
Published: (2023)
by: Kemertas, Mete, et al.
Published: (2023)
Press Start to Charge: Videogaming the Online Centralized Charging Scheduling Problem
by: Ghahtarani, Alireza, et al.
Published: (2026)
by: Ghahtarani, Alireza, et al.
Published: (2026)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
by: Farebrother, Jesse, et al.
Published: (2024)
by: Farebrother, Jesse, et al.
Published: (2024)
A Truncated Newton Method for Optimal Transport
by: Kemertas, Mete, et al.
Published: (2025)
by: Kemertas, Mete, et al.
Published: (2025)
SPEQ: Offline Stabilization Phases for Efficient Q-Learning in High Update-To-Data Ratio Reinforcement Learning
by: Romeo, Carlo, et al.
Published: (2025)
by: Romeo, Carlo, et al.
Published: (2025)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Is Value Learning Really the Main Bottleneck in Offline RL?
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
Transitive RL: Value Learning via Divide and Conquer
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Rényi Divergence Deep Mutual Learning
by: Huang, Weipeng, et al.
Published: (2022)
by: Huang, Weipeng, et al.
Published: (2022)
AFU: Actor-Free critic Updates in off-policy RL for continuous control
by: Perrin-Gilbert, Nicolas
Published: (2024)
by: Perrin-Gilbert, Nicolas
Published: (2024)
Influence functions and regularity tangents for efficient active learning
by: Eaton, Frederik
Published: (2024)
by: Eaton, Frederik
Published: (2024)
Improving Adversarial Transferability via Model Alignment
by: Ma, Avery, et al.
Published: (2023)
by: Ma, Avery, et al.
Published: (2023)
ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
by: Cherepanov, Egor, et al.
Published: (2025)
by: Cherepanov, Egor, et al.
Published: (2025)
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy Labels
by: Hu, Ruofan, et al.
Published: (2025)
by: Hu, Ruofan, et al.
Published: (2025)
Replicable Reinforcement Learning with Linear Function Approximation
by: Eaton, Eric, et al.
Published: (2025)
by: Eaton, Eric, et al.
Published: (2025)
Generalized Cauchy-Schwarz Divergence and Its Deep Learning Applications
by: Lu, Mingfei, et al.
Published: (2024)
by: Lu, Mingfei, et al.
Published: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
by: Dayal, Sarthak, et al.
Published: (2026)
by: Dayal, Sarthak, et al.
Published: (2026)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
by: Agrawalla, Bhavya, et al.
Published: (2025)
by: Agrawalla, Bhavya, et al.
Published: (2025)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
Similar Items
-
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
by: Voelcker, Claas A, et al.
Published: (2024) -
$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
by: Voelcker, Claas A, et al.
Published: (2023) -
When does Self-Prediction help? Understanding Auxiliary Tasks in Reinforcement Learning
by: Voelcker, Claas, et al.
Published: (2024) -
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025) -
Behavior-Consistent Deep Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2026)