MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Khadka, Barsat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CTS-Bench: Benchmarking Graph Coarsening Trade-offs for GNNs in Clock Tree Synthesis
von: Khadka, Barsat, et al.
Veröffentlicht: (2026)
von: Khadka, Barsat, et al.
Veröffentlicht: (2026)
Filter-then-Verify: A Multiphase GNN and ModernBERT Framework for Social Engineering Detection in Email Networks
von: Khadka, Barsat, et al.
Veröffentlicht: (2026)
von: Khadka, Barsat, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
Teaching Language Models Mechanistic Explainability Through MechSMILES
von: Neukomm, Théo A., et al.
Veröffentlicht: (2025)
von: Neukomm, Théo A., et al.
Veröffentlicht: (2025)
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
von: Hadad, Itamar, et al.
Veröffentlicht: (2026)
von: Hadad, Itamar, et al.
Veröffentlicht: (2026)
MechPert: Mechanistic Consensus as an Inductive Bias for Unseen Perturbation Prediction
von: Martell, Marc Boubnovski, et al.
Veröffentlicht: (2026)
von: Martell, Marc Boubnovski, et al.
Veröffentlicht: (2026)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
von: Nainani, Jatin
Veröffentlicht: (2024)
von: Nainani, Jatin
Veröffentlicht: (2024)
AutoResearch-RL: Perpetual Self-Evaluating Reinforcement Learning Agents for Autonomous Neural Architecture Discovery
von: Jain, Nilesh, et al.
Veröffentlicht: (2026)
von: Jain, Nilesh, et al.
Veröffentlicht: (2026)
Interpreting Reinforcement Learning Agents with Susceptibilities
von: Elliott, Chris, et al.
Veröffentlicht: (2026)
von: Elliott, Chris, et al.
Veröffentlicht: (2026)
Mechanistic Analysis of Circuit Preservation in Federated Learning
von: Haseeb, Muhammad, et al.
Veröffentlicht: (2025)
von: Haseeb, Muhammad, et al.
Veröffentlicht: (2025)
GenCircuit-RL: Reinforcement Learning from Hierarchical Verification for Genetic Circuit Design
von: Flynn, Noah
Veröffentlicht: (2026)
von: Flynn, Noah
Veröffentlicht: (2026)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
AstRL: Analog and Mixed-Signal Circuit Synthesis with Deep Reinforcement Learning
von: Guo, Felicia B., et al.
Veröffentlicht: (2026)
von: Guo, Felicia B., et al.
Veröffentlicht: (2026)
MechDetect: Detecting Data-Dependent Errors
von: Jung, Philipp, et al.
Veröffentlicht: (2025)
von: Jung, Philipp, et al.
Veröffentlicht: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
von: Kim, Geonhee, et al.
Veröffentlicht: (2024)
von: Kim, Geonhee, et al.
Veröffentlicht: (2024)
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
von: Xia, Peng, et al.
Veröffentlicht: (2026)
von: Xia, Peng, et al.
Veröffentlicht: (2026)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
von: Bai, Hao, et al.
Veröffentlicht: (2024)
von: Bai, Hao, et al.
Veröffentlicht: (2024)
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement Learning
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
SINDy-RL: Interpretable and Efficient Model-Based Reinforcement Learning
von: Zolman, Nicholas, et al.
Veröffentlicht: (2024)
von: Zolman, Nicholas, et al.
Veröffentlicht: (2024)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Compact Proofs of Model Performance via Mechanistic Interpretability
von: Gross, Jason, et al.
Veröffentlicht: (2024)
von: Gross, Jason, et al.
Veröffentlicht: (2024)
StackFeat RL: Reinforcement Learning over Iterative Dual Criterion Feature Selection for Stable Biomarker Discovery
von: Yermekov, A., et al.
Veröffentlicht: (2026)
von: Yermekov, A., et al.
Veröffentlicht: (2026)
DeepMech: A Machine Learning Framework for Chemical Reaction Mechanism Prediction
von: Das, Manajit, et al.
Veröffentlicht: (2025)
von: Das, Manajit, et al.
Veröffentlicht: (2025)
Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
von: Delfosse, Quentin, et al.
Veröffentlicht: (2024)
von: Delfosse, Quentin, et al.
Veröffentlicht: (2024)
PokeRL: Reinforcement Learning for Pokemon Red
von: Mudireddy, Dheeraj, et al.
Veröffentlicht: (2026)
von: Mudireddy, Dheeraj, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
Validating Mechanistic Interpretations: An Axiomatic Approach
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
von: Palumbo, Nils, et al.
Veröffentlicht: (2024)
On Mechanistic Circuits for Extractive Question-Answering
von: Basu, Samyadeep, et al.
Veröffentlicht: (2025)
von: Basu, Samyadeep, et al.
Veröffentlicht: (2025)
MechProNet: Machine Learning Prediction of Mechanical Properties in Metal Additive Manufacturing
von: Akbari, Parand, et al.
Veröffentlicht: (2022)
von: Akbari, Parand, et al.
Veröffentlicht: (2022)
Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits
von: Mahlau, Yannik, et al.
Veröffentlicht: (2025)
von: Mahlau, Yannik, et al.
Veröffentlicht: (2025)
MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG Discovery
von: Li, Dong, et al.
Veröffentlicht: (2026)
von: Li, Dong, et al.
Veröffentlicht: (2026)
RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
von: Lei, Kun, et al.
Veröffentlicht: (2025)
von: Lei, Kun, et al.
Veröffentlicht: (2025)
Separating Ansatz Discovery from Deployment on Larger Problems: Reinforcement Learning for Modular Circuit Design
von: Turati, Gloria, et al.
Veröffentlicht: (2025)
von: Turati, Gloria, et al.
Veröffentlicht: (2025)
Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
von: Masip, Sergi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CTS-Bench: Benchmarking Graph Coarsening Trade-offs for GNNs in Clock Tree Synthesis
von: Khadka, Barsat, et al.
Veröffentlicht: (2026) -
Filter-then-Verify: A Multiphase GNN and ModernBERT Framework for Social Engineering Detection in Email Networks
von: Khadka, Barsat, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024) -
Teaching Language Models Mechanistic Explainability Through MechSMILES
von: Neukomm, Théo A., et al.
Veröffentlicht: (2025) -
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
von: Hadad, Itamar, et al.
Veröffentlicht: (2026)