Interpreting Emergent Planning in Model-Free Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bush, Thomas, Chung, Stephen, Anwar, Usman, Garriga-Alonso, Adrià, Krueger, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026)
von: Chanin, David, et al.
Veröffentlicht: (2026)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
von: Arcuschin, Iván, et al.
Veröffentlicht: (2026)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2026)
Planning in a recurrent neural network that plays Sokoban
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
Learning to Forget using Hypernetworks
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
DiFR: Inference Verification Despite Nondeterminism
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Robust Model-Based Reinforcement Learning with an Adversarial Auxiliary Model
von: Herremans, Siemen, et al.
Veröffentlicht: (2024)
von: Herremans, Siemen, et al.
Veröffentlicht: (2024)
Hypothesis Testing the Circuit Hypothesis in LLMs
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
von: Xie, Sean, et al.
Veröffentlicht: (2022)
von: Xie, Sean, et al.
Veröffentlicht: (2022)
Prism: Policy Reuse via Interpretable Strategy Mapping in Reinforcement Learning
von: Pravetz, Thomas
Veröffentlicht: (2026)
von: Pravetz, Thomas
Veröffentlicht: (2026)
Surrogate Fitness Metrics for Interpretable Reinforcement Learning
von: Altmann, Philipp, et al.
Veröffentlicht: (2025)
von: Altmann, Philipp, et al.
Veröffentlicht: (2025)
Wavelet-Enhanced Neural ODE and Graph Attention for Interpretable Energy Forecasting
von: Joy, Usman Gani
Veröffentlicht: (2025)
von: Joy, Usman Gani
Veröffentlicht: (2025)
Demystifying MuZero Planning: Interpreting the Learned Model
von: Guei, Hung, et al.
Veröffentlicht: (2024)
von: Guei, Hung, et al.
Veröffentlicht: (2024)
Learning from Failures in Multi-Attempt Reinforcement Learning
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reinforcement Learning Algorithms for Autonomous Shipping
von: Lesy, Bavo, et al.
Veröffentlicht: (2024)
von: Lesy, Bavo, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning with Universal Horizon Models
von: Chung, Hojun, et al.
Veröffentlicht: (2026)
von: Chung, Hojun, et al.
Veröffentlicht: (2026)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
von: Eaton, Kenneth, et al.
Veröffentlicht: (2024)
von: Eaton, Kenneth, et al.
Veröffentlicht: (2024)
Parseval Regularization for Continual Reinforcement Learning
von: Chung, Wesley, et al.
Veröffentlicht: (2024)
von: Chung, Wesley, et al.
Veröffentlicht: (2024)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
von: Gupta, Rohan, et al.
Veröffentlicht: (2024)
Continual Reinforcement Learning by Planning with Online World Models
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Gradient Free Deep Reinforcement Learning With TabPFN
von: Schiff, David, et al.
Veröffentlicht: (2025)
von: Schiff, David, et al.
Veröffentlicht: (2025)
Towards General-Purpose Model-Free Reinforcement Learning
von: Fujimoto, Scott, et al.
Veröffentlicht: (2025)
von: Fujimoto, Scott, et al.
Veröffentlicht: (2025)
GOPlan: Goal-conditioned Offline Reinforcement Learning by Planning with Learned Models
von: Wang, Mianchu, et al.
Veröffentlicht: (2023)
von: Wang, Mianchu, et al.
Veröffentlicht: (2023)
Three Pathways to Neurosymbolic Reinforcement Learning with Interpretable Model and Policy Networks
von: Graf, Peter, et al.
Veröffentlicht: (2024)
von: Graf, Peter, et al.
Veröffentlicht: (2024)
Reward Model Ensembles Help Mitigate Overoptimization
von: Coste, Thomas, et al.
Veröffentlicht: (2023)
von: Coste, Thomas, et al.
Veröffentlicht: (2023)
Interpretable Deep Reinforcement Learning for Element-level Bridge Life-cycle Optimization
von: Moayyedi, Seyyed Amirhossein, et al.
Veröffentlicht: (2026)
von: Moayyedi, Seyyed Amirhossein, et al.
Veröffentlicht: (2026)
Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning
von: Son, Jaehyeon, et al.
Veröffentlicht: (2025)
von: Son, Jaehyeon, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Traffic Light Control in Intelligent Transportation Systems
von: Zhu, Ming, et al.
Veröffentlicht: (2023)
von: Zhu, Ming, et al.
Veröffentlicht: (2023)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
von: Wang, Yudan, et al.
Veröffentlicht: (2024)
von: Wang, Yudan, et al.
Veröffentlicht: (2024)
Label-Free Reinforcement Learning via Cross-Model Entropy
von: Gorbett, Matt, et al.
Veröffentlicht: (2026)
von: Gorbett, Matt, et al.
Veröffentlicht: (2026)
Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
von: Li, Zhaoxin, et al.
Veröffentlicht: (2025)
von: Li, Zhaoxin, et al.
Veröffentlicht: (2025)
Neural ODE and SDE Models for Adaptation and Planning in Model-Based Reinforcement Learning
von: Han, Chao, et al.
Veröffentlicht: (2026)
von: Han, Chao, et al.
Veröffentlicht: (2026)
Handling Delay in Real-Time Reinforcement Learning
von: Anokhin, Ivan, et al.
Veröffentlicht: (2025)
von: Anokhin, Ivan, et al.
Veröffentlicht: (2025)
Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
von: Kwa, Thomas, et al.
Veröffentlicht: (2024)
von: Kwa, Thomas, et al.
Veröffentlicht: (2024)
Physics-Inspired Interpretability Of Machine Learning Models
von: Niroomand, Maximilian P, et al.
Veröffentlicht: (2023)
von: Niroomand, Maximilian P, et al.
Veröffentlicht: (2023)
Interpretability by Design for Efficient Multi-Objective Reinforcement Learning
von: Xia, Qiyue, et al.
Veröffentlicht: (2025)
von: Xia, Qiyue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026) -
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025) -
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025) -
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)