Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
Fuente:
arXiv
Salvato in:
| Autori principali: | Hadad, Itamar, Katz, Guy, Bassan, Shahaf |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hard to Explain: On the Computational Hardness of In-Distribution Model Interpretation
di: Amir, Guy, et al.
Pubblicazione: (2024)
di: Amir, Guy, et al.
Pubblicazione: (2024)
Local vs. Global Interpretability: A Computational Complexity Perspective
di: Bassan, Shahaf, et al.
Pubblicazione: (2024)
di: Bassan, Shahaf, et al.
Pubblicazione: (2024)
What makes an Ensemble (Un) Interpretable?
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
di: Boetius, David, et al.
Pubblicazione: (2026)
di: Boetius, David, et al.
Pubblicazione: (2026)
Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
Provably Explaining Neural Additive Models
di: Bassan, Shahaf, et al.
Pubblicazione: (2026)
di: Bassan, Shahaf, et al.
Pubblicazione: (2026)
On the Computational Tractability of the (Many) Shapley Values
di: Marzouk, Reda, et al.
Pubblicazione: (2025)
di: Marzouk, Reda, et al.
Pubblicazione: (2025)
Explain Yourself, Briefly! Self-Explaining Neural Networks with Concise Sufficient Reasons
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
On Improving Deep Active Learning with Formal Verification
di: Spiegelman, Jonathan, et al.
Pubblicazione: (2025)
di: Spiegelman, Jonathan, et al.
Pubblicazione: (2025)
Marabou 2.0: A Versatile Formal Analyzer of Neural Networks
di: Wu, Haoze, et al.
Pubblicazione: (2024)
di: Wu, Haoze, et al.
Pubblicazione: (2024)
SHAP Meets Tensor Networks: Provably Tractable Explanations with Parallelism
di: Marzouk, Reda, et al.
Pubblicazione: (2025)
di: Marzouk, Reda, et al.
Pubblicazione: (2025)
Unifying Formal Explanations: A Complexity-Theoretic Perspective
di: Bassan, Shahaf, et al.
Pubblicazione: (2026)
di: Bassan, Shahaf, et al.
Pubblicazione: (2026)
Compact Proofs of Model Performance via Mechanistic Interpretability
di: Gross, Jason, et al.
Pubblicazione: (2024)
di: Gross, Jason, et al.
Pubblicazione: (2024)
Verifying the Generalization of Deep Learning to Out-of-Distribution Domains
di: Amir, Guy, et al.
Pubblicazione: (2024)
di: Amir, Guy, et al.
Pubblicazione: (2024)
Proof Minimization in Neural Network Verification
di: Isac, Omri, et al.
Pubblicazione: (2025)
di: Isac, Omri, et al.
Pubblicazione: (2025)
PICID: Proof-Driven Clause Learning in Neural Network Verification
di: Isac, Omri, et al.
Pubblicazione: (2025)
di: Isac, Omri, et al.
Pubblicazione: (2025)
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
di: Aljaafari, Nura, et al.
Pubblicazione: (2026)
di: Aljaafari, Nura, et al.
Pubblicazione: (2026)
Additive Models Explained: A Computational Complexity Approach
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)
A Compositional Atlas for Algebraic Circuits
di: Wang, Benjie, et al.
Pubblicazione: (2024)
di: Wang, Benjie, et al.
Pubblicazione: (2024)
Shield Synthesis for LTL Modulo Theories
di: Rodriguez, Andoni, et al.
Pubblicazione: (2024)
di: Rodriguez, Andoni, et al.
Pubblicazione: (2024)
Pseudo-Formalization for Automatic Proof Verification
di: Barkallah, Slim, et al.
Pubblicazione: (2026)
di: Barkallah, Slim, et al.
Pubblicazione: (2026)
Formalized Hopfield Networks and Boltzmann Machines
di: Cipollina, Matteo, et al.
Pubblicazione: (2025)
di: Cipollina, Matteo, et al.
Pubblicazione: (2025)
Towards a Certified Proof Checker for Deep Neural Network Verification
di: Desmartin, Remi, et al.
Pubblicazione: (2023)
di: Desmartin, Remi, et al.
Pubblicazione: (2023)
Towards Verifiable Transformers: Solver-Checkable Circuit Explanations
di: Somani, Neel
Pubblicazione: (2026)
di: Somani, Neel
Pubblicazione: (2026)
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
di: Yousefzadeh, Roozbeh, et al.
Pubblicazione: (2025)
di: Yousefzadeh, Roozbeh, et al.
Pubblicazione: (2025)
Provable Preimage Under-Approximation for Neural Networks (Full Version)
di: Zhang, Xiyue, et al.
Pubblicazione: (2023)
di: Zhang, Xiyue, et al.
Pubblicazione: (2023)
Floating-Point Neural Networks Are Provably Robust Universal Approximators
di: Hwang, Geonho, et al.
Pubblicazione: (2025)
di: Hwang, Geonho, et al.
Pubblicazione: (2025)
Formal Explanations for Neuro-Symbolic AI
di: Paul, Sushmita, et al.
Pubblicazione: (2024)
di: Paul, Sushmita, et al.
Pubblicazione: (2024)
Learning Temporal Logic Predicates from Data with Statistical Guarantees
di: Soroka, Emi, et al.
Pubblicazione: (2024)
di: Soroka, Emi, et al.
Pubblicazione: (2024)
FAME: Formal Abstract Minimal Explanation for Neural Networks
di: Boumazouza, Ryma, et al.
Pubblicazione: (2026)
di: Boumazouza, Ryma, et al.
Pubblicazione: (2026)
What are the Right Symmetries for Formal Theorem Proving?
di: Olejniczak, Krzysztof, et al.
Pubblicazione: (2026)
di: Olejniczak, Krzysztof, et al.
Pubblicazione: (2026)
Formal Mathematical Reasoning: A New Frontier in AI
di: Yang, Kaiyu, et al.
Pubblicazione: (2024)
di: Yang, Kaiyu, et al.
Pubblicazione: (2024)
Solving Formal Math Problems by Decomposition and Iterative Reflection
di: Zhou, Yichi, et al.
Pubblicazione: (2025)
di: Zhou, Yichi, et al.
Pubblicazione: (2025)
LeanAgent: Lifelong Learning for Formal Theorem Proving
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2024)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2024)
Circuit Representations of Random Forests with Applications to XAI
di: Ji, Chunxi, et al.
Pubblicazione: (2026)
di: Ji, Chunxi, et al.
Pubblicazione: (2026)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
How (and when) can you fit examples to logic-based hypothesis classes over infinite structures?
di: Benedikt, Michael, et al.
Pubblicazione: (2026)
di: Benedikt, Michael, et al.
Pubblicazione: (2026)
Logical GANs: Adversarial Learning through Ehrenfeucht Fraisse Games
di: Mannucci, Mirco A.
Pubblicazione: (2025)
di: Mannucci, Mirco A.
Pubblicazione: (2025)
From learnable objects to learnable random objects
di: Anderson, Aaron, et al.
Pubblicazione: (2025)
di: Anderson, Aaron, et al.
Pubblicazione: (2025)
Programs as Singularities
di: Murfet, Daniel, et al.
Pubblicazione: (2025)
di: Murfet, Daniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Hard to Explain: On the Computational Hardness of In-Distribution Model Interpretation
di: Amir, Guy, et al.
Pubblicazione: (2024) -
Local vs. Global Interpretability: A Computational Complexity Perspective
di: Bassan, Shahaf, et al.
Pubblicazione: (2024) -
What makes an Ensemble (Un) Interpretable?
di: Bassan, Shahaf, et al.
Pubblicazione: (2025) -
Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
di: Boetius, David, et al.
Pubblicazione: (2026) -
Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations
di: Bassan, Shahaf, et al.
Pubblicazione: (2025)