InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Rohan, Arcuschin, Iván, Kwa, Thomas, Garriga-Alonso, Adrià |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
di: Kwa, Thomas, et al.
Pubblicazione: (2024)
di: Kwa, Thomas, et al.
Pubblicazione: (2024)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
di: Chanin, David, et al.
Pubblicazione: (2026)
di: Chanin, David, et al.
Pubblicazione: (2026)
Adversarial Circuit Evaluation
di: de Bos, Niels uit, et al.
Pubblicazione: (2024)
di: de Bos, Niels uit, et al.
Pubblicazione: (2024)
Investigating the Indirect Object Identification circuit in Mamba
di: Ensign, Danielle, et al.
Pubblicazione: (2024)
di: Ensign, Danielle, et al.
Pubblicazione: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2025)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
di: Bush, Thomas, et al.
Pubblicazione: (2025)
di: Bush, Thomas, et al.
Pubblicazione: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Compact Proofs of Model Performance via Mechanistic Interpretability
di: Gross, Jason, et al.
Pubblicazione: (2024)
di: Gross, Jason, et al.
Pubblicazione: (2024)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Open Problems in Mechanistic Interpretability
di: Sharkey, Lee, et al.
Pubblicazione: (2025)
di: Sharkey, Lee, et al.
Pubblicazione: (2025)
Automatically Finding Reward Model Biases
di: Wang, Atticus, et al.
Pubblicazione: (2026)
di: Wang, Atticus, et al.
Pubblicazione: (2026)
Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop
di: Kohler, Hector, et al.
Pubblicazione: (2024)
di: Kohler, Hector, et al.
Pubblicazione: (2024)
Mechanistic Interpretability for Transformer-based Time Series Classification
di: Kalnāre, Matīss, et al.
Pubblicazione: (2025)
di: Kalnāre, Matīss, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
From Mechanistic to Compositional Interpretability
di: Gauderis, Ward, et al.
Pubblicazione: (2026)
di: Gauderis, Ward, et al.
Pubblicazione: (2026)
DiFR: Inference Verification Despite Nondeterminism
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Analyzing the Generalization and Reliability of Steering Vectors
di: Tan, Daniel, et al.
Pubblicazione: (2024)
di: Tan, Daniel, et al.
Pubblicazione: (2024)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
di: Dumas, Clément
Pubblicazione: (2025)
di: Dumas, Clément
Pubblicazione: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
Inference-Time Toxicity Mitigation in Protein Language Models
di: Burda, Manuel Fernández, et al.
Pubblicazione: (2026)
di: Burda, Manuel Fernández, et al.
Pubblicazione: (2026)
Base Models Know How to Reason, Thinking Models Learn When
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
di: El, Batu, et al.
Pubblicazione: (2025)
di: El, Batu, et al.
Pubblicazione: (2025)
Exemplar Partitioning for Mechanistic Interpretability
di: Rumbelow, Jessica
Pubblicazione: (2026)
di: Rumbelow, Jessica
Pubblicazione: (2026)
Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering
di: Sprejer, Eitan, et al.
Pubblicazione: (2026)
di: Sprejer, Eitan, et al.
Pubblicazione: (2026)
SplInterp: Improving our Understanding and Training of Sparse Autoencoders
di: Budd, Jeremy, et al.
Pubblicazione: (2025)
di: Budd, Jeremy, et al.
Pubblicazione: (2025)
Planning in a recurrent neural network that plays Sokoban
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2024)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Reinforcement Learning Agents
di: Trim, Tristan, et al.
Pubblicazione: (2024)
di: Trim, Tristan, et al.
Pubblicazione: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
di: Palumbo, Nils, et al.
Pubblicazione: (2024)
di: Palumbo, Nils, et al.
Pubblicazione: (2024)
Mechanistic Interpretability for Neural TSP Solvers
di: Narad, Reuben, et al.
Pubblicazione: (2025)
di: Narad, Reuben, et al.
Pubblicazione: (2025)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
di: Sutter, Denis, et al.
Pubblicazione: (2025)
di: Sutter, Denis, et al.
Pubblicazione: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
di: Gonzalez, ML Nissen, et al.
Pubblicazione: (2026)
di: Gonzalez, ML Nissen, et al.
Pubblicazione: (2026)
On the Limits of Interpretable Machine Learning in Quintic Root Classification
di: Thomas, Rohan, et al.
Pubblicazione: (2026)
di: Thomas, Rohan, et al.
Pubblicazione: (2026)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
Contrast Sets for Evaluating Language-Guided Robot Policies
di: Anwar, Abrar, et al.
Pubblicazione: (2024)
di: Anwar, Abrar, et al.
Pubblicazione: (2024)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
di: Meek, Austin, et al.
Pubblicazione: (2025)
di: Meek, Austin, et al.
Pubblicazione: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
di: Nainani, Jatin, et al.
Pubblicazione: (2024)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
di: Kwa, Thomas, et al.
Pubblicazione: (2024) -
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
di: Arcuschin, Iván, et al.
Pubblicazione: (2026) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
di: Chanin, David, et al.
Pubblicazione: (2026) -
Adversarial Circuit Evaluation
di: de Bos, Niels uit, et al.
Pubblicazione: (2024) -
Investigating the Indirect Object Identification circuit in Mamba
di: Ensign, Danielle, et al.
Pubblicazione: (2024)