Interpretability Can Be Actionable
Fuente:
arXiv
Saved in:
| Main Authors: | Orgad, Hadas, Barez, Fazl, Haklay, Tal, Lee, Isabelle, Mosbach, Marius, Reusch, Anja, Saphra, Naomi, Wallace, Byron, Wiegreffe, Sarah, Wong, Eric, Tenney, Ian, Geva, Mor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026)
by: Rai, Daking, et al.
Published: (2026)
Mechanistic?
by: Saphra, Naomi, et al.
Published: (2024)
by: Saphra, Naomi, et al.
Published: (2024)
Position-aware Automatic Circuit Discovery
by: Haklay, Tal, et al.
Published: (2025)
by: Haklay, Tal, et al.
Published: (2025)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
by: Mosbach, Marius, et al.
Published: (2024)
by: Mosbach, Marius, et al.
Published: (2024)
Precise In-Parameter Concept Erasure in Large Language Models
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
The Hidden Space of Transformer Language Adapters
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Query Circuits: Explaining How Language Models Answer User Prompts
by: Wu, Tung-Yu, et al.
Published: (2025)
by: Wu, Tung-Yu, et al.
Published: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Rigorous Interpretation Is a Form of Evaluation
by: Lee, Isabelle, et al.
Published: (2026)
by: Lee, Isabelle, et al.
Published: (2026)
Can Interpretation Predict Behavior on Unseen Data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Rethinking AI Cultural Alignment
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
A Graph-based Framework for Coverage Analysis in Autonomous Driving
by: Muehlenstädt, Thomas, et al.
Published: (2026)
by: Muehlenstädt, Thomas, et al.
Published: (2026)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
by: Gottesman, Daniela, et al.
Published: (2025)
by: Gottesman, Daniela, et al.
Published: (2025)
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024)
by: Fung, Samy Wu, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Do Activation Verbalization Methods Convey Privileged Information?
by: Li, Millicent, et al.
Published: (2025)
by: Li, Millicent, et al.
Published: (2025)
Contents of hydrocarbons in waters and bottom sediments of the southwestern Amur Bay (Sea of Japan) in April 2005
by: Nemirovskaya, Inna A
Published: (2007)
by: Nemirovskaya, Inna A
Published: (2007)
Physical oceanography during Bjarni Saemundsson cruise B07/99
by: VEINS Members, et al.
Published: (2011)
by: VEINS Members, et al.
Published: (2011)
Physical oceanography during Håkon Mosby cruise HM07/1 to the North Sea
by: IYFS/IBTS
Published: (2010)
by: IYFS/IBTS
Published: (2010)
Soil bulk density and soil depth from on-site observations in the North-Western Kurdistan region, Iraq
by: Bellat, Mathias, et al.
Published: (2024)
by: Bellat, Mathias, et al.
Published: (2024)
Physical oceanography during Thalassa cruise Thalassa07/1 to the North Sea
by: IYFS/IBTS
Published: (2010)
by: IYFS/IBTS
Published: (2010)
Hydrography measured in lakes during the Saskylakh expedition in 2007
by: Herzschuh, Ulrike
Published: (2016)
by: Herzschuh, Ulrike
Published: (2016)
Token Taxes: mitigating AGI's economic risks
by: Irwin, Lucas, et al.
Published: (2026)
by: Irwin, Lucas, et al.
Published: (2026)
Large Language Models Relearn Removed Concepts
by: Lo, Michelle, et al.
Published: (2024)
by: Lo, Michelle, et al.
Published: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Hydrochemistry measured on water bottle samples during Bjarni Saemundsson cruise B07/99
by: VEINS Members, et al.
Published: (2011)
by: VEINS Members, et al.
Published: (2011)
Age and chemical composition of magmatic rocks from the Belkovsky Island, New Siberian Islands
by: Kuzmichev, A B, et al.
Published: (2008)
by: Kuzmichev, A B, et al.
Published: (2008)
OZ-7 Lake Eyre region in South Australia and along the Darling River up to Bourke in New South Wales sampling trip and analysis: strontium, rubidium and neodymium isotopes
by: De Deckker, Patrick
Published: (2019)
by: De Deckker, Patrick
Published: (2019)
Physical oceanography during Dana II cruise Dana07/1 to the North Sea
by: IYFS/IBTS
Published: (2010)
by: IYFS/IBTS
Published: (2010)
Similar Items
-
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026) -
Mechanistic?
by: Saphra, Naomi, et al.
Published: (2024) -
Position-aware Automatic Circuit Discovery
by: Haklay, Tal, et al.
Published: (2025) -
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024) -
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
by: Mosbach, Marius, et al.
Published: (2024)