Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching
Fuente:
arXiv
Saved in:
| Main Author: | Bahador, Nooshin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Semi-Supervised Anomaly Detection Pipeline for SOZ Localization Using Ictal-Related Chirp
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Interpretable Concept-Based Memory Reasoning
by: Debot, David, et al.
Published: (2024)
by: Debot, David, et al.
Published: (2024)
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
by: Gao, Zijun, et al.
Published: (2025)
by: Gao, Zijun, et al.
Published: (2025)
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
by: Poonia, Ansh, et al.
Published: (2025)
by: Poonia, Ansh, et al.
Published: (2025)
Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
PatchDecomp: Interpretable Patch-Based Time Series Forecasting
by: Tomioka, Hiroki, et al.
Published: (2026)
by: Tomioka, Hiroki, et al.
Published: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
Interpretable Hierarchical Concept Reasoning through Attention-Guided Graph Learning
by: Debot, David, et al.
Published: (2025)
by: Debot, David, et al.
Published: (2025)
Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
by: Chorna, Sofiia, et al.
Published: (2025)
by: Chorna, Sofiia, et al.
Published: (2025)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Interpretable Neural-Symbolic Concept Reasoning
by: Barbiero, Pietro, et al.
Published: (2023)
by: Barbiero, Pietro, et al.
Published: (2023)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Learning Concept Bottleneck Models from Mechanistic Explanations
by: De Santis, Antonio, et al.
Published: (2026)
by: De Santis, Antonio, et al.
Published: (2026)
SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability
by: Zhang, Hanwei, et al.
Published: (2026)
by: Zhang, Hanwei, et al.
Published: (2026)
Classifying Clinical Outcome of Epilepsy Patients with Ictal Chirp Embeddings
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
by: Lin, Zihao, et al.
Published: (2025)
by: Lin, Zihao, et al.
Published: (2025)
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
Measuring the Depth of LLM Unlearning via Activation Patching
by: Lee, Jaeung, et al.
Published: (2026)
by: Lee, Jaeung, et al.
Published: (2026)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
by: Ren, Zirui, et al.
Published: (2026)
by: Ren, Zirui, et al.
Published: (2026)
A Mechanistic Analysis of Looped Reasoning Language Models
by: Blayney, Hugh, et al.
Published: (2026)
by: Blayney, Hugh, et al.
Published: (2026)
From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Mind The Gap: Quantifying Mechanistic Gaps in Algorithmic Reasoning via Neural Compilation
by: Saldyt, Lucas, et al.
Published: (2025)
by: Saldyt, Lucas, et al.
Published: (2025)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
by: Maghsoudi, Maryam, et al.
Published: (2026)
by: Maghsoudi, Maryam, et al.
Published: (2026)
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
by: Helff, Lukas, et al.
Published: (2025)
by: Helff, Lukas, et al.
Published: (2025)
Towards Autonomous Mechanistic Reasoning in Virtual Cells
by: Jang, Yunhui, et al.
Published: (2026)
by: Jang, Yunhui, et al.
Published: (2026)
Leakage and Interpretability in Concept-Based Models
by: Parisini, Enrico, et al.
Published: (2025)
by: Parisini, Enrico, et al.
Published: (2025)
Hierarchical Concept-based Interpretable Models
by: Hill, Oscar, et al.
Published: (2026)
by: Hill, Oscar, et al.
Published: (2026)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
by: Liang, Jia, et al.
Published: (2026)
by: Liang, Jia, et al.
Published: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Revealing Combinatorial Reasoning of GNNs via Graph Concept Bottleneck Layer
by: Niu, Yue, et al.
Published: (2026)
by: Niu, Yue, et al.
Published: (2026)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
Similar Items
-
Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
by: Bahador, Nooshin
Published: (2025) -
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study
by: Bahador, Nooshin, et al.
Published: (2025) -
Semi-Supervised Anomaly Detection Pipeline for SOZ Localization Using Ictal-Related Chirp
by: Bahador, Nooshin, et al.
Published: (2025) -
Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework
by: Bahador, Nooshin
Published: (2025) -
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)