Mechanistic Interpretability of Reinforcement Learning Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Trim, Tristan, Grayston, Triston |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
by: Khadka, Barsat
Published: (2026)
by: Khadka, Barsat
Published: (2026)
Interpreting Reinforcement Learning Agents with Susceptibilities
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
by: Miller, Ryan J., et al.
Published: (2025)
by: Miller, Ryan J., et al.
Published: (2025)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
Distributed Area Coverage with High Altitude Balloons Using Multi-Agent Reinforcement Learning
by: Haroon, Adam, et al.
Published: (2025)
by: Haroon, Adam, et al.
Published: (2025)
Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
by: Delfosse, Quentin, et al.
Published: (2024)
by: Delfosse, Quentin, et al.
Published: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
by: Palumbo, Nils, et al.
Published: (2024)
by: Palumbo, Nils, et al.
Published: (2024)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
by: Masip, Sergi, et al.
Published: (2026)
by: Masip, Sergi, et al.
Published: (2026)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Geospatial Mechanistic Interpretability of Large Language Models
by: De Sabbata, Stef, et al.
Published: (2025)
by: De Sabbata, Stef, et al.
Published: (2025)
Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: a Short Survey
by: Colas, Cédric, et al.
Published: (2020)
by: Colas, Cédric, et al.
Published: (2020)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
by: Yan, John, et al.
Published: (2026)
by: Yan, John, et al.
Published: (2026)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024)
by: Li, Jason
Published: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
by: Torre, Elia, et al.
Published: (2025)
by: Torre, Elia, et al.
Published: (2025)
Interpretable Learning Dynamics in Unsupervised Reinforcement Learning
by: Pandey, Shashwat
Published: (2025)
by: Pandey, Shashwat
Published: (2025)
Interpretable Failure Analysis in Multi-Agent Reinforcement Learning Systems
by: Shefin, Risal Shahriar, et al.
Published: (2026)
by: Shefin, Risal Shahriar, et al.
Published: (2026)
Interpreting Agent Behaviors in Reinforcement-Learning-Based Cyber-Battle Simulation Platforms
by: Claypoole, Jared, et al.
Published: (2025)
by: Claypoole, Jared, et al.
Published: (2025)
The Impact of Quantization on Large Reasoning Model Reinforcement Learning
by: Kumar, Medha, et al.
Published: (2025)
by: Kumar, Medha, et al.
Published: (2025)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
by: Tomilin, Tristan, et al.
Published: (2025)
by: Tomilin, Tristan, et al.
Published: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
by: Erdogan, Ege, et al.
Published: (2025)
by: Erdogan, Ege, et al.
Published: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
by: Gonzalez, ML Nissen, et al.
Published: (2026)
by: Gonzalez, ML Nissen, et al.
Published: (2026)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
by: Gupta, Rohan, et al.
Published: (2024)
by: Gupta, Rohan, et al.
Published: (2024)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
by: Sutter, Denis, et al.
Published: (2025)
by: Sutter, Denis, et al.
Published: (2025)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
by: Kowalska, Bianka, et al.
Published: (2025)
by: Kowalska, Bianka, et al.
Published: (2025)
Group-Agent Reinforcement Learning with Heterogeneous Agents
by: Wu, Kaiyue, et al.
Published: (2025)
by: Wu, Kaiyue, et al.
Published: (2025)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Optimizing Interpretable Decision Tree Policies for Reinforcement Learning
by: Vos, Daniël, et al.
Published: (2024)
by: Vos, Daniël, et al.
Published: (2024)
Safety-Oriented Pruning and Interpretation of Reinforcement Learning Policies
by: Gross, Dennis, et al.
Published: (2024)
by: Gross, Dennis, et al.
Published: (2024)
Similar Items
-
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
by: Khadka, Barsat
Published: (2026) -
Interpreting Reinforcement Learning Agents with Susceptibilities
by: Elliott, Chris, et al.
Published: (2026) -
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
by: Miller, Ryan J., et al.
Published: (2025) -
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026) -
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)