Validating Mechanistic Interpretations: An Axiomatic Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Palumbo, Nils, Mangal, Ravi, Wang, Zifan, Vijayakumar, Saranya, Pasareanu, Corina S., Jha, Somesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Concept-Based Abductive and Contrastive Explanations for Behaviors of Vision Models
von: Canizales, Ronaldo, et al.
Veröffentlicht: (2026)
von: Canizales, Ronaldo, et al.
Veröffentlicht: (2026)
Scenario-based Compositional Verification of Autonomous Systems with Neural Perception
von: Watson, Christopher, et al.
Veröffentlicht: (2025)
von: Watson, Christopher, et al.
Veröffentlicht: (2025)
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
von: Hu, Boyue Caroline, et al.
Veröffentlicht: (2025)
von: Hu, Boyue Caroline, et al.
Veröffentlicht: (2025)
How Not to Detect Prompt Injections with an LLM
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025)
On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection
von: Hannan, Md Abdul, et al.
Veröffentlicht: (2025)
von: Hannan, Md Abdul, et al.
Veröffentlicht: (2025)
Concept-based Analysis of Neural Networks via Vision-Language Models
von: Mangal, Ravi, et al.
Veröffentlicht: (2024)
von: Mangal, Ravi, et al.
Veröffentlicht: (2024)
Conformal Safety Shielding for Imperfect-Perception Agents
von: Scarbro, William, et al.
Veröffentlicht: (2025)
von: Scarbro, William, et al.
Veröffentlicht: (2025)
Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and Rejection
von: Palumbo, Nils, et al.
Veröffentlicht: (2023)
von: Palumbo, Nils, et al.
Veröffentlicht: (2023)
Prophecy: Inferring Formal Properties from Neuron Activations
von: Gopinath, Divya, et al.
Veröffentlicht: (2025)
von: Gopinath, Divya, et al.
Veröffentlicht: (2025)
Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
von: Melo, Rui, et al.
Veröffentlicht: (2025)
von: Melo, Rui, et al.
Veröffentlicht: (2025)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2026)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2026)
SLVR: Securely Leveraging Client Validation for Robust Federated Learning
von: Choi, Jihye, et al.
Veröffentlicht: (2025)
von: Choi, Jihye, et al.
Veröffentlicht: (2025)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
Probably Approximately Correct Causal Discovery
von: Wei, Mian, et al.
Veröffentlicht: (2025)
von: Wei, Mian, et al.
Veröffentlicht: (2025)
A Mixture of Linear Corrections Generates Secure Code
von: Yu, Weichen, et al.
Veröffentlicht: (2025)
von: Yu, Weichen, et al.
Veröffentlicht: (2025)
MALADE: Orchestration of LLM-powered Agents with Retrieval Augmented Generation for Pharmacovigilance
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
An Axiomatic Approach to Model-Agnostic Concept Explanations
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
An Axiomatic Approach to Loss Aggregation and an Adapted Aggregating Algorithm
von: Pacheco, Armando J. Cabrera, et al.
Veröffentlicht: (2024)
von: Pacheco, Armando J. Cabrera, et al.
Veröffentlicht: (2024)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Interpretable Hybrid Deep Q-Learning Framework for IoT-Based Food Spoilage Prediction with Synthetic Data Generation and Hardware Validation
von: Singh, Isshaan, et al.
Veröffentlicht: (2025)
von: Singh, Isshaan, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
von: Trim, Tristan, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability for Neural TSP Solvers
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
von: Narad, Reuben, et al.
Veröffentlicht: (2025)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
Dependency-Aware Privacy for Multi-turn Agents
von: Anshumaan, Divyam, et al.
Veröffentlicht: (2026)
von: Anshumaan, Divyam, et al.
Veröffentlicht: (2026)
Verifier-Guided Code Translation via Meta-Step Decoding
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
ASPEST: Bridging the Gap Between Active Learning and Selective Prediction
von: Chen, Jiefeng, et al.
Veröffentlicht: (2023)
von: Chen, Jiefeng, et al.
Veröffentlicht: (2023)
Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Geospatial Mechanistic Interpretability of Large Language Models
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
von: De Sabbata, Stef, et al.
Veröffentlicht: (2025)
Learning Interpretable Differentiable Logic Networks
von: Yue, Chang, et al.
Veröffentlicht: (2024)
von: Yue, Chang, et al.
Veröffentlicht: (2024)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2025)
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
von: He, Jesse, et al.
Veröffentlicht: (2026)
von: He, Jesse, et al.
Veröffentlicht: (2026)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
von: Torre, Elia, et al.
Veröffentlicht: (2025)
von: Torre, Elia, et al.
Veröffentlicht: (2025)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
von: Miller, Ryan J., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Concept-Based Abductive and Contrastive Explanations for Behaviors of Vision Models
von: Canizales, Ronaldo, et al.
Veröffentlicht: (2026) -
Scenario-based Compositional Verification of Autonomous Systems with Neural Perception
von: Watson, Christopher, et al.
Veröffentlicht: (2025) -
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
von: Hu, Boyue Caroline, et al.
Veröffentlicht: (2025) -
How Not to Detect Prompt Injections with an LLM
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2025) -
On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection
von: Hannan, Md Abdul, et al.
Veröffentlicht: (2025)