Interpreting Reinforcement Learning Agents with Susceptibilities
Fuente:
arXiv
Saved in:
| Main Authors: | Elliott, Chris, Urdshals, Einar, Quarel, David, Murfet, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Structure Development in List-Sorting Transformers
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Linear Response Estimators for Singular Statistical Models
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Structural Inference: Interpreting Small Language Models with Susceptibilities
by: Baker, Garrett, et al.
Published: (2025)
by: Baker, Garrett, et al.
Published: (2025)
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
by: Benito-Rodriguez, Éloïse, et al.
Published: (2025)
by: Benito-Rodriguez, Éloïse, et al.
Published: (2025)
Patterning: The Dual of Interpretability
by: Wang, George, et al.
Published: (2026)
by: Wang, George, et al.
Published: (2026)
Modes of Sequence Models and Learning Coefficients
by: Chen, Zhongtian, et al.
Published: (2025)
by: Chen, Zhongtian, et al.
Published: (2025)
Towards Spectroscopy: Susceptibility Clusters in Language Models
by: Gordon, Andrew, et al.
Published: (2026)
by: Gordon, Andrew, et al.
Published: (2026)
Programs as Singularities
by: Murfet, Daniel, et al.
Published: (2025)
by: Murfet, Daniel, et al.
Published: (2025)
Dark Matter-induced electron excitations in silicon and germanium with Deep Learning
by: Catena, Riccardo, et al.
Published: (2024)
by: Catena, Riccardo, et al.
Published: (2024)
Mechanistic Interpretability of Reinforcement Learning Agents
by: Trim, Tristan, et al.
Published: (2024)
by: Trim, Tristan, et al.
Published: (2024)
Embryology of a Language Model
by: Wang, George, et al.
Published: (2025)
by: Wang, George, et al.
Published: (2025)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
by: Fillingham, Sean P., et al.
Published: (2025)
by: Fillingham, Sean P., et al.
Published: (2025)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
by: Carroll, Liam, et al.
Published: (2025)
by: Carroll, Liam, et al.
Published: (2025)
Symbolic State Partitioning for Reinforcement Learning
by: Ghaffari, Mohsen, et al.
Published: (2024)
by: Ghaffari, Mohsen, et al.
Published: (2024)
Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
by: Delfosse, Quentin, et al.
Published: (2024)
by: Delfosse, Quentin, et al.
Published: (2024)
Optimizing Interpretable Decision Tree Policies for Reinforcement Learning
by: Vos, Daniël, et al.
Published: (2024)
by: Vos, Daniël, et al.
Published: (2024)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)
by: Wang, George, et al.
Published: (2024)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
by: Khadka, Barsat
Published: (2026)
by: Khadka, Barsat
Published: (2026)
Factored Value Functions for Graph-Based Multi-Agent Reinforcement Learning
by: Rashwan, Ahmed, et al.
Published: (2026)
by: Rashwan, Ahmed, et al.
Published: (2026)
Deep Reinforcement Learning for Autonomous Cyber Defence: A Survey
by: Palmer, Gregory, et al.
Published: (2023)
by: Palmer, Gregory, et al.
Published: (2023)
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
by: Yan, John, et al.
Published: (2026)
by: Yan, John, et al.
Published: (2026)
Interpretable Learning Dynamics in Unsupervised Reinforcement Learning
by: Pandey, Shashwat
Published: (2025)
by: Pandey, Shashwat
Published: (2025)
Learning When to Switch: Adaptive Policy Selection via Reinforcement Learning
by: Tava, Chris
Published: (2025)
by: Tava, Chris
Published: (2025)
Interpretable Failure Analysis in Multi-Agent Reinforcement Learning Systems
by: Shefin, Risal Shahriar, et al.
Published: (2026)
by: Shefin, Risal Shahriar, et al.
Published: (2026)
Interpreting Agent Behaviors in Reinforcement-Learning-Based Cyber-Battle Simulation Platforms
by: Claypoole, Jared, et al.
Published: (2025)
by: Claypoole, Jared, et al.
Published: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
by: Bush, Thomas, et al.
Published: (2025)
by: Bush, Thomas, et al.
Published: (2025)
Group-Agent Reinforcement Learning with Heterogeneous Agents
by: Wu, Kaiyue, et al.
Published: (2025)
by: Wu, Kaiyue, et al.
Published: (2025)
Principal Prototype Analysis on Manifold for Interpretable Reinforcement Learning
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
Safety-Oriented Pruning and Interpretation of Reinforcement Learning Policies
by: Gross, Dennis, et al.
Published: (2024)
by: Gross, Dennis, et al.
Published: (2024)
Interpretability of Statistical, Machine Learning, and Deep Learning Models for Landslide Susceptibility Mapping in Three Gorges Reservoir Area
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
Reinforcement Learning applied to Insurance Portfolio Pursuit
by: Young, Edward James, et al.
Published: (2024)
by: Young, Edward James, et al.
Published: (2024)
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
by: Lau, Edmund, et al.
Published: (2023)
by: Lau, Edmund, et al.
Published: (2023)
Probabilistic Constrained Reinforcement Learning with Formal Interpretability
by: Wang, Yanran, et al.
Published: (2023)
by: Wang, Yanran, et al.
Published: (2023)
Multi-Agent Reinforcement Learning in Intelligent Transportation Systems: A Comprehensive Survey
by: Donatus, Rexcharles, et al.
Published: (2025)
by: Donatus, Rexcharles, et al.
Published: (2025)
Interpretable Deep Reinforcement Learning for Element-level Bridge Life-cycle Optimization
by: Moayyedi, Seyyed Amirhossein, et al.
Published: (2026)
by: Moayyedi, Seyyed Amirhossein, et al.
Published: (2026)
Upside-Down Reinforcement Learning for More Interpretable Optimal Control
by: Cardenas-Cartagena, Juan, et al.
Published: (2024)
by: Cardenas-Cartagena, Juan, et al.
Published: (2024)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
by: Xie, Sean, et al.
Published: (2022)
by: Xie, Sean, et al.
Published: (2022)
Similar Items
-
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
by: Elliott, Chris, et al.
Published: (2026) -
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
by: Elliott, Chris, et al.
Published: (2026) -
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
by: Urdshals, Einar, et al.
Published: (2025) -
Structure Development in List-Sorting Transformers
by: Urdshals, Einar, et al.
Published: (2025) -
Linear Response Estimators for Singular Statistical Models
by: Elliott, Chris, et al.
Published: (2026)