Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Zhengxuan, Geiger, Atticus, Icard, Thomas, Potts, Christopher, Goodman, Noah D. |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
par: Geiger, Atticus, et autres
Publié: (2023)
par: Geiger, Atticus, et autres
Publié: (2023)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
par: Wu, Zhengxuan, et autres
Publié: (2024)
par: Wu, Zhengxuan, et autres
Publié: (2024)
How Causal Abstraction Underpins Computational Explanation
par: Geiger, Atticus, et autres
Publié: (2025)
par: Geiger, Atticus, et autres
Publié: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
par: Huang, Jing, et autres
Publié: (2024)
par: Huang, Jing, et autres
Publié: (2024)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
par: Wu, Zhengxuan, et autres
Publié: (2024)
par: Wu, Zhengxuan, et autres
Publié: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
par: Sun, Jiuding, et autres
Publié: (2025)
par: Sun, Jiuding, et autres
Publié: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
par: Puyin, Li, et autres
Publié: (2026)
par: Puyin, Li, et autres
Publié: (2026)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
par: Geiger, Atticus, et autres
Publié: (2023)
par: Geiger, Atticus, et autres
Publié: (2023)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
par: Wu, Zhengxuan, et autres
Publié: (2023)
par: Wu, Zhengxuan, et autres
Publié: (2023)
ReFT: Representation Finetuning for Language Models
par: Wu, Zhengxuan, et autres
Publié: (2024)
par: Wu, Zhengxuan, et autres
Publié: (2024)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
par: Huang, Jing, et autres
Publié: (2025)
par: Huang, Jing, et autres
Publié: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
par: Shafran, Or, et autres
Publié: (2025)
par: Shafran, Or, et autres
Publié: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
par: Wu, Zhengxuan, et autres
Publié: (2025)
par: Wu, Zhengxuan, et autres
Publié: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
par: Sun, Jiuding, et autres
Publié: (2025)
par: Sun, Jiuding, et autres
Publié: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
par: She, Jingyuan Selena, et autres
Publié: (2023)
par: She, Jingyuan Selena, et autres
Publié: (2023)
Updating CLIP to Prefer Descriptions Over Captions
par: Zur, Amir, et autres
Publié: (2024)
par: Zur, Amir, et autres
Publié: (2024)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
Activation Steering via Generative Causal Mediation
par: Sankaranarayanan, Aruna, et autres
Publié: (2026)
par: Sankaranarayanan, Aruna, et autres
Publié: (2026)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
par: Zhong, Zexuan, et autres
Publié: (2023)
par: Zhong, Zexuan, et autres
Publié: (2023)
Improved Representation Steering for Language Models
par: Wu, Zhengxuan, et autres
Publié: (2025)
par: Wu, Zhengxuan, et autres
Publié: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
par: Wu, Yiwei, et autres
Publié: (2025)
par: Wu, Yiwei, et autres
Publié: (2025)
Bayesian scaling laws for in-context learning
par: Arora, Aryaman, et autres
Publié: (2024)
par: Arora, Aryaman, et autres
Publié: (2024)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
par: Csordás, Róbert, et autres
Publié: (2024)
par: Csordás, Róbert, et autres
Publié: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
par: Zur, Amir, et autres
Publié: (2025)
par: Zur, Amir, et autres
Publié: (2025)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
par: Mao, Jiayuan, et autres
Publié: (2023)
par: Mao, Jiayuan, et autres
Publié: (2023)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
par: Shafran, Or, et autres
Publié: (2026)
par: Shafran, Or, et autres
Publié: (2026)
Endless Terminals: Scaling RL Environments for Terminal Agents
par: Gandhi, Kanishk, et autres
Publié: (2026)
par: Gandhi, Kanishk, et autres
Publié: (2026)
Learning to Simulate Human Dialogue
par: Gandhi, Kanishk, et autres
Publié: (2026)
par: Gandhi, Kanishk, et autres
Publié: (2026)
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
par: Boguraev, Sasha, et autres
Publié: (2025)
par: Boguraev, Sasha, et autres
Publié: (2025)
Oolong: Investigating What Makes Transfer Learning Hard with Controlled Studies
par: Wu, Zhengxuan, et autres
Publié: (2022)
par: Wu, Zhengxuan, et autres
Publié: (2022)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
par: Mishra, Shubhra, et autres
Publié: (2024)
par: Mishra, Shubhra, et autres
Publié: (2024)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
par: Yu, Qinan, et autres
Publié: (2026)
par: Yu, Qinan, et autres
Publié: (2026)
PreFT: Prefill-only finetuning for efficient inference
par: Lanpouthakoun, Andrew, et autres
Publié: (2026)
par: Lanpouthakoun, Andrew, et autres
Publié: (2026)
Learning to Compress Prompts with Gist Tokens
par: Mu, Jesse, et autres
Publié: (2023)
par: Mu, Jesse, et autres
Publié: (2023)
Scaling up the think-aloud method
par: Wurgaft, Daniel, et autres
Publié: (2025)
par: Wurgaft, Daniel, et autres
Publié: (2025)
Interleaving Logic and Counting
par: van Benthem, Johan, et autres
Publié: (2025)
par: van Benthem, Johan, et autres
Publié: (2025)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
par: He-Yueya, Joy, et autres
Publié: (2024)
par: He-Yueya, Joy, et autres
Publié: (2024)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
par: Arora, Aryaman, et autres
Publié: (2024)
par: Arora, Aryaman, et autres
Publié: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
par: Chaudhary, Maheep, et autres
Publié: (2024)
par: Chaudhary, Maheep, et autres
Publié: (2024)
Documents similaires
-
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
par: Geiger, Atticus, et autres
Publié: (2023) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
par: Wu, Zhengxuan, et autres
Publié: (2024) -
How Causal Abstraction Underpins Computational Explanation
par: Geiger, Atticus, et autres
Publié: (2025) -
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
par: Huang, Jing, et autres
Publié: (2024) -
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
par: Wu, Zhengxuan, et autres
Publié: (2024)