Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Geiger, Atticus, Wu, Zhengxuan, Potts, Christopher, Icard, Thomas, Goodman, Noah D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
von: Geiger, Atticus, et al.
Veröffentlicht: (2023)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
von: Puyin, Li, et al.
Veröffentlicht: (2026)
von: Puyin, Li, et al.
Veröffentlicht: (2026)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Combining Causal Models for More Accurate Abstractions of Neural Networks
von: Pîslar, Theodora-Mara, et al.
Veröffentlicht: (2025)
von: Pîslar, Theodora-Mara, et al.
Veröffentlicht: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Emergent Symbol-like Number Variables in Artificial Neural Networks
von: Grant, Satchel, et al.
Veröffentlicht: (2025)
von: Grant, Satchel, et al.
Veröffentlicht: (2025)
Modeling Discrimination with Causal Abstraction
von: Mossé, Milan, et al.
Veröffentlicht: (2025)
von: Mossé, Milan, et al.
Veröffentlicht: (2025)
On Probabilistic and Causal Reasoning with Summation Operators
von: Ibeling, Duligur, et al.
Veröffentlicht: (2024)
von: Ibeling, Duligur, et al.
Veröffentlicht: (2024)
Bayesian scaling laws for in-context learning
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
Automatically Finding Reward Model Biases
von: Wang, Atticus, et al.
Veröffentlicht: (2026)
von: Wang, Atticus, et al.
Veröffentlicht: (2026)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
von: Zur, Amir, et al.
Veröffentlicht: (2025)
von: Zur, Amir, et al.
Veröffentlicht: (2025)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
von: Mao, Jiayuan, et al.
Veröffentlicht: (2023)
von: Mao, Jiayuan, et al.
Veröffentlicht: (2023)
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
von: Boguraev, Sasha, et al.
Veröffentlicht: (2025)
von: Boguraev, Sasha, et al.
Veröffentlicht: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
von: Mishra, Shubhra, et al.
Veröffentlicht: (2024)
von: Mishra, Shubhra, et al.
Veröffentlicht: (2024)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
PreFT: Prefill-only finetuning for efficient inference
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
von: Lanpouthakoun, Andrew, et al.
Veröffentlicht: (2026)
A Flexible Method for Behaviorally Measuring Alignment Between Human and Artificial Intelligence Using Representational Similarity Analysis
von: Ogg, Mattson, et al.
Veröffentlicht: (2024)
von: Ogg, Mattson, et al.
Veröffentlicht: (2024)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
von: Conan, Jean-Baptiste A.
Veröffentlicht: (2025)
von: Conan, Jean-Baptiste A.
Veröffentlicht: (2025)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
From Creation to Curriculum: Examining the role of generative AI in Arts Universities
von: Sims, Atticus
Veröffentlicht: (2024)
von: Sims, Atticus
Veröffentlicht: (2024)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
CARL: Causality-guided Architecture Representation Learning for an Interpretable Performance Predictor
von: Ji, Han, et al.
Veröffentlicht: (2025)
von: Ji, Han, et al.
Veröffentlicht: (2025)
Large Language Model Reasoning Failures
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
Learning Formal Mathematics From Intrinsic Motivation
von: Poesia, Gabriel, et al.
Veröffentlicht: (2024)
von: Poesia, Gabriel, et al.
Veröffentlicht: (2024)
Graph Similarity Computation via Interpretable Neural Node Alignment
von: Wang, Jingjing, et al.
Veröffentlicht: (2024)
von: Wang, Jingjing, et al.
Veröffentlicht: (2024)
Resource Rational Contractualism Should Guide AI Alignment
von: Levine, Sydney, et al.
Veröffentlicht: (2025)
von: Levine, Sydney, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024) -
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
von: Geiger, Atticus, et al.
Veröffentlicht: (2023) -
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025) -
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)