MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Jesse, Jenne, Helen, Vargas, Max, Brown, Davis, Mishne, Gal, Wang, Yusu, Kvinge, Henry |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
Machine Learning meets Algebraic Combinatorics: A Suite of Datasets Capturing Research-level Conjecturing Ability in Pure Mathematics
di: Chau, Herman, et al.
Pubblicazione: (2025)
di: Chau, Herman, et al.
Pubblicazione: (2025)
Can Neural Networks Learn Small Algebraic Worlds? An Investigation Into the Group-theoretic Structures Learned By Narrow Models Trained To Predict Group Operations
di: Kvinge, Henry, et al.
Pubblicazione: (2026)
di: Kvinge, Henry, et al.
Pubblicazione: (2026)
Explaining GNN Explanations with Edge Gradients
di: He, Jesse, et al.
Pubblicazione: (2025)
di: He, Jesse, et al.
Pubblicazione: (2025)
Machines and Mathematical Mutations: Using GNNs to Characterize Quiver Mutation Classes
di: He, Jesse, et al.
Pubblicazione: (2024)
di: He, Jesse, et al.
Pubblicazione: (2024)
Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
Model editing for distribution shifts in uranium oxide morphological analysis
di: Brown, Davis, et al.
Pubblicazione: (2024)
di: Brown, Davis, et al.
Pubblicazione: (2024)
Elucidating Flow Matching ODE Dynamics with Respect to Data Geometries and Denoisers
di: Wan, Zhengchao, et al.
Pubblicazione: (2024)
di: Wan, Zhengchao, et al.
Pubblicazione: (2024)
Mind The Gap: Quantifying Mechanistic Gaps in Algorithmic Reasoning via Neural Compilation
di: Saldyt, Lucas, et al.
Pubblicazione: (2025)
di: Saldyt, Lucas, et al.
Pubblicazione: (2025)
The Numerical Stability of Hyperbolic Representation Learning
di: Mishne, Gal, et al.
Pubblicazione: (2022)
di: Mishne, Gal, et al.
Pubblicazione: (2022)
Neural Algorithmic Reasoning with Multiple Correct Solutions
di: Kujawa, Zeno, et al.
Pubblicazione: (2024)
di: Kujawa, Zeno, et al.
Pubblicazione: (2024)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
General and Efficient Steering of Unconditional Diffusion
di: Wang, Qingsong, et al.
Pubblicazione: (2026)
di: Wang, Qingsong, et al.
Pubblicazione: (2026)
What do Geometric Hallucination Detection Metrics Actually Measure?
di: Yeats, Eric, et al.
Pubblicazione: (2026)
di: Yeats, Eric, et al.
Pubblicazione: (2026)
Comparing Graph Transformers via Positional Encodings
di: Black, Mitchell, et al.
Pubblicazione: (2024)
di: Black, Mitchell, et al.
Pubblicazione: (2024)
Challenges in Mechanistically Interpreting Model Representations
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
Interpretable Neural Networks with Random Constructive Algorithm
di: Nan, Jing, et al.
Pubblicazione: (2023)
di: Nan, Jing, et al.
Pubblicazione: (2023)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching
di: Bahador, Nooshin
Pubblicazione: (2025)
di: Bahador, Nooshin
Pubblicazione: (2025)
Deep Neural Network Benchmarks for Selective Classification
di: Pugnana, Andrea, et al.
Pubblicazione: (2024)
di: Pugnana, Andrea, et al.
Pubblicazione: (2024)
Open-Book Neural Algorithmic Reasoning
di: Li, Hefei, et al.
Pubblicazione: (2024)
di: Li, Hefei, et al.
Pubblicazione: (2024)
Recurrent Aggregators in Neural Algorithmic Reasoning
di: Xu, Kaijia, et al.
Pubblicazione: (2024)
di: Xu, Kaijia, et al.
Pubblicazione: (2024)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
di: Dumas, Clément
Pubblicazione: (2025)
di: Dumas, Clément
Pubblicazione: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
di: Yan, Ge, et al.
Pubblicazione: (2025)
di: Yan, Ge, et al.
Pubblicazione: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
di: Kalnāre, Matīss, et al.
Pubblicazione: (2025)
di: Kalnāre, Matīss, et al.
Pubblicazione: (2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
di: Lin, Zihao, et al.
Pubblicazione: (2025)
di: Lin, Zihao, et al.
Pubblicazione: (2025)
Neural Interpretable Reasoning
di: Barbiero, Pietro, et al.
Pubblicazione: (2025)
di: Barbiero, Pietro, et al.
Pubblicazione: (2025)
PUZZLES: A Benchmark for Neural Algorithmic Reasoning
di: Estermann, Benjamin, et al.
Pubblicazione: (2024)
di: Estermann, Benjamin, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
di: Nadaf, Mohammed Suhail B
Pubblicazione: (2026)
di: Nadaf, Mohammed Suhail B
Pubblicazione: (2026)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
di: El, Batu, et al.
Pubblicazione: (2025)
di: El, Batu, et al.
Pubblicazione: (2025)
Towards Autonomous Mechanistic Reasoning in Virtual Cells
di: Jang, Yunhui, et al.
Pubblicazione: (2026)
di: Jang, Yunhui, et al.
Pubblicazione: (2026)
Richer Representations for Neural Algorithmic Reasoning via Auxiliary Reconstruction
di: Huang, Jiafu, et al.
Pubblicazione: (2026)
di: Huang, Jiafu, et al.
Pubblicazione: (2026)
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
di: Ren, Zirui, et al.
Pubblicazione: (2026)
di: Ren, Zirui, et al.
Pubblicazione: (2026)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Which Algorithms Can Graph Neural Networks Learn?
di: Wittig, Solveig, et al.
Pubblicazione: (2026)
di: Wittig, Solveig, et al.
Pubblicazione: (2026)
A Mechanistic Analysis of Looped Reasoning Language Models
di: Blayney, Hugh, et al.
Pubblicazione: (2026)
di: Blayney, Hugh, et al.
Pubblicazione: (2026)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
di: Jacobi, Jonathan, et al.
Pubblicazione: (2025)
di: Jacobi, Jonathan, et al.
Pubblicazione: (2025)
Neural Probabilistic Circuits: Enabling Compositional and Interpretable Predictions through Logical Reasoning
di: Chen, Weixin, et al.
Pubblicazione: (2025)
di: Chen, Weixin, et al.
Pubblicazione: (2025)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
di: Brown, Davis, et al.
Pubblicazione: (2025) -
Machine Learning meets Algebraic Combinatorics: A Suite of Datasets Capturing Research-level Conjecturing Ability in Pure Mathematics
di: Chau, Herman, et al.
Pubblicazione: (2025) -
Can Neural Networks Learn Small Algebraic Worlds? An Investigation Into the Group-theoretic Structures Learned By Narrow Models Trained To Predict Group Operations
di: Kvinge, Henry, et al.
Pubblicazione: (2026) -
Explaining GNN Explanations with Edge Gradients
di: He, Jesse, et al.
Pubblicazione: (2025) -
Machines and Mathematical Mutations: Using GNNs to Characterize Quiver Mutation Classes
di: He, Jesse, et al.
Pubblicazione: (2024)