MIB: A Mechanistic Interpretability Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Mueller, Aaron, Geiger, Atticus, Wiegreffe, Sarah, Arad, Dana, Arcuschin, Iván, Belfki, Adam, Chan, Yik Siu, Fiotto-Kaufman, Jaden, Haklay, Tal, Hanna, Michael, Huang, Jing, Gupta, Rohan, Nikankin, Yaniv, Orgad, Hadas, Prakash, Nikhil, Reusch, Anja, Sankaranarayanan, Aruna, Shao, Shun, Stolfo, Alessandro, Tutek, Martin, Zur, Amir, Bau, David, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Position-aware Automatic Circuit Discovery
di: Haklay, Tal, et al.
Pubblicazione: (2025)
di: Haklay, Tal, et al.
Pubblicazione: (2025)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023)
di: Arad, Dana, et al.
Pubblicazione: (2023)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
Reasoning Models Know What's Important, and Encode It in Their Activations
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Activation Steering via Generative Causal Mediation
di: Sankaranarayanan, Aruna, et al.
Pubblicazione: (2026)
di: Sankaranarayanan, Aruna, et al.
Pubblicazione: (2026)
Unified Concept Editing in Diffusion Models
di: Gandikota, Rohit, et al.
Pubblicazione: (2023)
di: Gandikota, Rohit, et al.
Pubblicazione: (2023)
Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models
di: Pomerants, Gal, et al.
Pubblicazione: (2026)
di: Pomerants, Gal, et al.
Pubblicazione: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
Reverse-Engineering the Retrieval Process in GenIR Models
di: Reusch, Anja, et al.
Pubblicazione: (2025)
di: Reusch, Anja, et al.
Pubblicazione: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Interpretability Can Be Actionable
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
di: Toker, Michael, et al.
Pubblicazione: (2025)
di: Toker, Michael, et al.
Pubblicazione: (2025)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
di: Tutek, Martin, et al.
Pubblicazione: (2025)
di: Tutek, Martin, et al.
Pubblicazione: (2025)
Language Models use Lookbacks to Track Beliefs
di: Prakash, Nikhil, et al.
Pubblicazione: (2025)
di: Prakash, Nikhil, et al.
Pubblicazione: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Automatically Finding Reward Model Biases
di: Wang, Atticus, et al.
Pubblicazione: (2026)
di: Wang, Atticus, et al.
Pubblicazione: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
Linearity of Relation Decoding in Transformer Language Models
di: Hernandez, Evan, et al.
Pubblicazione: (2023)
di: Hernandez, Evan, et al.
Pubblicazione: (2023)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
di: Zur, Amir, et al.
Pubblicazione: (2025)
di: Zur, Amir, et al.
Pubblicazione: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2026)
di: Simhi, Adi, et al.
Pubblicazione: (2026)
Inside-Out: Hidden Factual Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Updating CLIP to Prefer Descriptions Over Captions
di: Zur, Amir, et al.
Pubblicazione: (2024)
di: Zur, Amir, et al.
Pubblicazione: (2024)
Confidence Regulation Neurons in Language Models
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
Agents of Chaos
di: Shapira, Natalie, et al.
Pubblicazione: (2026)
di: Shapira, Natalie, et al.
Pubblicazione: (2026)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
di: Mueller, Aaron, et al.
Pubblicazione: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
di: Shafran, Or, et al.
Pubblicazione: (2025)
di: Shafran, Or, et al.
Pubblicazione: (2025)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
di: Rudman, William, et al.
Pubblicazione: (2026)
di: Rudman, William, et al.
Pubblicazione: (2026)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
di: Marks, Samuel, et al.
Pubblicazione: (2024)
di: Marks, Samuel, et al.
Pubblicazione: (2024)
How Causal Abstraction Underpins Computational Explanation
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Position-aware Automatic Circuit Discovery
di: Haklay, Tal, et al.
Pubblicazione: (2025) -
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025) -
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023) -
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025) -
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)