Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
Fuente:
arXiv
Saved in:
| Main Author: | Mueller, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective for Molecular Property Prediction
by: He, Yinhan, et al.
Published: (2024)
by: He, Yinhan, et al.
Published: (2024)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
by: Kim, Hazel H.
Published: (2024)
by: Kim, Hazel H.
Published: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
Counterfactual Generation with Identifiability Guarantees
by: Yan, Hanqi, et al.
Published: (2024)
by: Yan, Hanqi, et al.
Published: (2024)
Aligning (Medical) LLMs for (Counterfactual) Fairness
by: Poulain, Raphael, et al.
Published: (2024)
by: Poulain, Raphael, et al.
Published: (2024)
AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment
by: Ivanova, Anastasiia, et al.
Published: (2025)
by: Ivanova, Anastasiia, et al.
Published: (2025)
Compared to What? Baselines and Metrics for Counterfactual Prompting
by: Yang, Zihao, et al.
Published: (2026)
by: Yang, Zihao, et al.
Published: (2026)
Neural Network Verification is a Programming Language Challenge
by: Cordeiro, Lucas C., et al.
Published: (2025)
by: Cordeiro, Lucas C., et al.
Published: (2025)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
iFlip: Iterative Feedback-driven Counterfactual Example Refinement
by: Wang, Yilong, et al.
Published: (2026)
by: Wang, Yilong, et al.
Published: (2026)
Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction
by: Goh, Hui Wen, et al.
Published: (2026)
by: Goh, Hui Wen, et al.
Published: (2026)
A Practical Method for Generating String Counterfactuals
by: Avitan, Matan, et al.
Published: (2024)
by: Avitan, Matan, et al.
Published: (2024)
Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model
by: Zhang, Yizhou, et al.
Published: (2023)
by: Zhang, Yizhou, et al.
Published: (2023)
Function Vectors in Large Language Models
by: Todd, Eric, et al.
Published: (2023)
by: Todd, Eric, et al.
Published: (2023)
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Iterative Counterfactual Data Augmentation
by: Plyler, Mitchell, et al.
Published: (2025)
by: Plyler, Mitchell, et al.
Published: (2025)
Similar Phrases for Cause of Actions of Civil Cases
by: Huang, Ho-Chien, et al.
Published: (2024)
by: Huang, Ho-Chien, et al.
Published: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
by: Joishy, Anish R, et al.
Published: (2025)
by: Joishy, Anish R, et al.
Published: (2025)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
by: Yoon, Youngsik, et al.
Published: (2026)
by: Yoon, Youngsik, et al.
Published: (2026)
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
by: Dotsinski, Asen, et al.
Published: (2025)
by: Dotsinski, Asen, et al.
Published: (2025)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
AxBERT: An Interpretable Chinese Spelling Correction Method Driven by Associative Knowledge Network
by: Wang, Fanyu, et al.
Published: (2025)
by: Wang, Fanyu, et al.
Published: (2025)
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Structural Pruning of Pre-trained Language Models via Neural Architecture Search
by: Klein, Aaron, et al.
Published: (2024)
by: Klein, Aaron, et al.
Published: (2024)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Explaining Text Classifiers with Counterfactual Representations
by: Lemberger, Pirmin, et al.
Published: (2024)
by: Lemberger, Pirmin, et al.
Published: (2024)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
by: Zhang, Zhexin, et al.
Published: (2026)
by: Zhang, Zhexin, et al.
Published: (2026)
Graph Neural Networks on Discriminative Graphs of Words
by: Abbahaddou, Yassine, et al.
Published: (2024)
by: Abbahaddou, Yassine, et al.
Published: (2024)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
Convolutional Neural Networks for Toxic Comment Classification
by: Georgakopoulos, Spiros V., et al.
Published: (2018)
by: Georgakopoulos, Spiros V., et al.
Published: (2018)
Article Classification with Graph Neural Networks and Multigraphs
by: Ly, Khang, et al.
Published: (2023)
by: Ly, Khang, et al.
Published: (2023)
A Survey : Neural Networks for AMR-to-Text
by: Hao, Hongyu, et al.
Published: (2022)
by: Hao, Hongyu, et al.
Published: (2022)
Graph Neural Network and NER-Based Text Summarization
by: Khan, Imaad Zaffar, et al.
Published: (2024)
by: Khan, Imaad Zaffar, et al.
Published: (2024)
Predicting Question Quality on StackOverflow with Neural Networks
by: Al-Ramahi, Mohammad, et al.
Published: (2024)
by: Al-Ramahi, Mohammad, et al.
Published: (2024)
Similar Items
-
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026) -
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024) -
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025) -
Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective for Molecular Property Prediction
by: He, Yinhan, et al.
Published: (2024) -
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
by: Kim, Hazel H.
Published: (2024)