BlackBoxToBlueprint: Extracting Interpretable Logic from Legacy Systems using Reinforcement Learning and Counterfactual Analysis
Fuente:
arXiv
Saved in:
| Main Author: | Rathore, Vidhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Manifold Fairness: Enforcing Geometric Invariance in Representation Learning
by: Rathore, Vidhi
Published: (2026)
by: Rathore, Vidhi
Published: (2026)
Benchmarking Instance-Centric Counterfactual Algorithms for XAI: From White Box to Black Box
by: Moreira, Catarina, et al.
Published: (2022)
by: Moreira, Catarina, et al.
Published: (2022)
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
by: Huang, Tairan, et al.
Published: (2025)
by: Huang, Tairan, et al.
Published: (2025)
Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding
by: Bethell, Daniel, et al.
Published: (2024)
by: Bethell, Daniel, et al.
Published: (2024)
Is Prior-Free Black-Box Non-Stationary Reinforcement Learning Feasible?
by: Gerogiannis, Argyrios, et al.
Published: (2024)
by: Gerogiannis, Argyrios, et al.
Published: (2024)
Cyclic Counterfactuals under Shift-Scale Interventions
by: Saha, Saptarshi, et al.
Published: (2025)
by: Saha, Saptarshi, et al.
Published: (2025)
Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement Learning
by: Miao, Ruiyao, et al.
Published: (2025)
by: Miao, Ruiyao, et al.
Published: (2025)
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
by: Moon, Jihoon
Published: (2025)
by: Moon, Jihoon
Published: (2025)
Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Counterfactual Explanations for Continuous Action Reinforcement Learning
by: Dong, Shuyang, et al.
Published: (2025)
by: Dong, Shuyang, et al.
Published: (2025)
Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models
by: Joshi, Aniruddha, et al.
Published: (2025)
by: Joshi, Aniruddha, et al.
Published: (2025)
Locally Interpretable Individualized Treatment Rules for Black-Box Decision Models
by: Charvadeh, Yasin Khadem, et al.
Published: (2026)
by: Charvadeh, Yasin Khadem, et al.
Published: (2026)
Reinforced In-Context Black-Box Optimization
by: Song, Lei, et al.
Published: (2024)
by: Song, Lei, et al.
Published: (2024)
Boosting deep Reinforcement Learning using pretraining with Logical Options
by: Ye, Zihan, et al.
Published: (2026)
by: Ye, Zihan, et al.
Published: (2026)
Benchmarking Counterfactual Interpretability in Deep Learning Models for Time Series Classification
by: Kan, Ziwen, et al.
Published: (2024)
by: Kan, Ziwen, et al.
Published: (2024)
Interpretable Failure Analysis in Multi-Agent Reinforcement Learning Systems
by: Shefin, Risal Shahriar, et al.
Published: (2026)
by: Shefin, Risal Shahriar, et al.
Published: (2026)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
by: Dagdanov, Resul, et al.
Published: (2022)
by: Dagdanov, Resul, et al.
Published: (2022)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
by: Chuck, Caleb, et al.
Published: (2025)
by: Chuck, Caleb, et al.
Published: (2025)
Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning
by: Vaskov, Sean, et al.
Published: (2024)
by: Vaskov, Sean, et al.
Published: (2024)
Extracting PAC Decision Trees from Black Box Binary Classifiers: The Gender Bias Case Study on BERT-based Language Models
by: Ozaki, Ana, et al.
Published: (2024)
by: Ozaki, Ana, et al.
Published: (2024)
Comparing Post-Hoc Explainable AI Methods for Interpreting Black-Box EEG Models in Depression Detection
by: Šarčević, Antonia, et al.
Published: (2026)
by: Šarčević, Antonia, et al.
Published: (2026)
LiBOG: Lifelong Learning for Black-Box Optimizer Generation
by: Pei, Jiyuan, et al.
Published: (2025)
by: Pei, Jiyuan, et al.
Published: (2025)
Explaining the Behavior of Black-Box Prediction Algorithms with Causal Learning
by: Sani, Numair, et al.
Published: (2020)
by: Sani, Numair, et al.
Published: (2020)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Learning Surrogates for Offline Black-Box Optimization via Gradient Matching
by: Hoang, Minh, et al.
Published: (2025)
by: Hoang, Minh, et al.
Published: (2025)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Bounding-Box Inference for Error-Aware Model-Based Reinforcement Learning
by: Talvitie, Erin J., et al.
Published: (2024)
by: Talvitie, Erin J., et al.
Published: (2024)
Optimistic Gradient Learning with Hessian Corrections for High-Dimensional Black-Box Optimization
by: Kfir, Yedidya, et al.
Published: (2025)
by: Kfir, Yedidya, et al.
Published: (2025)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
by: Bhambri, Siddhant, et al.
Published: (2024)
by: Bhambri, Siddhant, et al.
Published: (2024)
In Search of Trees: Decision-Tree Policy Synthesis for Black-Box Systems via Search
by: Demirović, Emir, et al.
Published: (2024)
by: Demirović, Emir, et al.
Published: (2024)
In-Context Black-Box Optimization with Unreliable Feedback
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
by: Xie, Sean, et al.
Published: (2022)
by: Xie, Sean, et al.
Published: (2022)
Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning
by: Bozkurt, Alper Kamil, et al.
Published: (2019)
by: Bozkurt, Alper Kamil, et al.
Published: (2019)
Surrogate Fitness Metrics for Interpretable Reinforcement Learning
by: Altmann, Philipp, et al.
Published: (2025)
by: Altmann, Philipp, et al.
Published: (2025)
Geometric Learning in Black-Box Optimization: A GNN Framework for Algorithm Performance Prediction
by: Kostovska, Ana, et al.
Published: (2025)
by: Kostovska, Ana, et al.
Published: (2025)
Similar Items
-
Causal Manifold Fairness: Enforcing Geometric Invariance in Representation Learning
by: Rathore, Vidhi
Published: (2026) -
Benchmarking Instance-Centric Counterfactual Algorithms for XAI: From White Box to Black Box
by: Moreira, Catarina, et al.
Published: (2022) -
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
by: Liu, Junhao, et al.
Published: (2025) -
SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
by: Huang, Tairan, et al.
Published: (2025) -
Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding
by: Bethell, Daniel, et al.
Published: (2024)