Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Kowalska, Bianka, Kwaśnicka, Halina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
by: Moon, Jihoon
Published: (2025)
by: Moon, Jihoon
Published: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Unraveling the Black Box of Neural Networks: A Dynamic Extremum Mapper
by: Chen, Shengjian
Published: (2025)
by: Chen, Shengjian
Published: (2025)
Integrating White and Black Box Techniques for Interpretable Machine Learning
by: Vernon, Eric M., et al.
Published: (2024)
by: Vernon, Eric M., et al.
Published: (2024)
Know2Vec: A Black-Box Proxy for Neural Network Retrieval
by: Shang, Zhuoyi, et al.
Published: (2024)
by: Shang, Zhuoyi, et al.
Published: (2024)
Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
by: Du, Mengyao, et al.
Published: (2025)
by: Du, Mengyao, et al.
Published: (2025)
La productivité dans les économies planifiées d'Europe orientale
by: Halina Kowalska
Published: (1956)
by: Halina Kowalska
Published: (1956)
La productividad en las economías planificadas de Europa oriental
by: Halina Kowalska
Published: (1956)
by: Halina Kowalska
Published: (1956)
Productivity in the planned economies of Eastern Europe
by: Halina Kowalska
Published: (1956)
by: Halina Kowalska
Published: (1956)
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
Interpretable Neural Networks with Random Constructive Algorithm
by: Nan, Jing, et al.
Published: (2023)
by: Nan, Jing, et al.
Published: (2023)
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
by: Yu, Shixing, et al.
Published: (2026)
by: Yu, Shixing, et al.
Published: (2026)
Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual Understanding
by: Xiong, Nuoye, et al.
Published: (2025)
by: Xiong, Nuoye, et al.
Published: (2025)
midr: Learning from Black-Box Models by Maximum Interpretation Decomposition
by: Asashiba, Ryoichi, et al.
Published: (2025)
by: Asashiba, Ryoichi, et al.
Published: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
by: Gonzalez, ML Nissen, et al.
Published: (2026)
by: Gonzalez, ML Nissen, et al.
Published: (2026)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Benchmarking Instance-Centric Counterfactual Algorithms for XAI: From White Box to Black Box
by: Moreira, Catarina, et al.
Published: (2022)
by: Moreira, Catarina, et al.
Published: (2022)
Beyond the Black Box: Interpretability of LLMs in Finance
by: Tatsat, Hariom, et al.
Published: (2025)
by: Tatsat, Hariom, et al.
Published: (2025)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
by: Chang, Shuochen, et al.
Published: (2026)
by: Chang, Shuochen, et al.
Published: (2026)
Concept Learning in the Wild: Towards Algorithmic Understanding of Neural Networks
by: Shoham, Elad, et al.
Published: (2024)
by: Shoham, Elad, et al.
Published: (2024)
Extracting Explanations, Justification, and Uncertainty from Black-Box Deep Neural Networks
by: Ardis, Paul, et al.
Published: (2024)
by: Ardis, Paul, et al.
Published: (2024)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Explaining the Behavior of Black-Box Prediction Algorithms with Causal Learning
by: Sani, Numair, et al.
Published: (2020)
by: Sani, Numair, et al.
Published: (2020)
Black-Box Forgetting
by: Kuwana, Yusuke, et al.
Published: (2024)
by: Kuwana, Yusuke, et al.
Published: (2024)
Unpacking the Black Box: Regulating Algorithmic Decisions
by: Blattner, Laura, et al.
Published: (2021)
by: Blattner, Laura, et al.
Published: (2021)
Smoothing the Black-Box: Signed-Distance Supervision for Black-Box Model Copying
by: Jiménez, Rubén, et al.
Published: (2026)
by: Jiménez, Rubén, et al.
Published: (2026)
Locally Interpretable Individualized Treatment Rules for Black-Box Decision Models
by: Charvadeh, Yasin Khadem, et al.
Published: (2026)
by: Charvadeh, Yasin Khadem, et al.
Published: (2026)
Mechanistic Interpretability of Reinforcement Learning Agents
by: Trim, Tristan, et al.
Published: (2024)
by: Trim, Tristan, et al.
Published: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
by: Palumbo, Nils, et al.
Published: (2024)
by: Palumbo, Nils, et al.
Published: (2024)
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
How Interpretable Are Interpretable Graph Neural Networks?
by: Chen, Yongqiang, et al.
Published: (2024)
by: Chen, Yongqiang, et al.
Published: (2024)
A Black-box Attack on Neural Networks Based on Swarm Evolutionary Algorithm
by: Liu, Xiaolei, et al.
Published: (2019)
by: Liu, Xiaolei, et al.
Published: (2019)
Black-Box Anomaly Attribution
by: Idé, Tsuyoshi, et al.
Published: (2023)
by: Idé, Tsuyoshi, et al.
Published: (2023)
Geospatial Mechanistic Interpretability of Large Language Models
by: De Sabbata, Stef, et al.
Published: (2025)
by: De Sabbata, Stef, et al.
Published: (2025)
Similar Items
-
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
by: Moon, Jihoon
Published: (2025) -
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026) -
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026) -
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025) -
Unraveling the Black Box of Neural Networks: A Dynamic Extremum Mapper
by: Chen, Shengjian
Published: (2025)